返回博客列表
OBPAI Geek Team2026/8/436 阅读

2026 LLM Comprehensive Capability Leaderboard

Based on the Arena 2026-08-01 and Artificial Analysis public data from recent months, we assign S/A/B grades across coding, long-context, Chinese/open-source, and Agent scenarios, and include a draggable LLM leaderboard template for custom revisions.

Large Language ModelsLLM LeaderboardArena RankingsModel SelectionAI
2026 LLM Comprehensive Capability Leaderboard - 天梯图与数据榜单

How the ranking works

Data as of 2026-08-05Scenarios — coding, long-form writing, Chinese/open-source, and agent — are graded S/A/B, based mainly on the Arena (arena.ai) August 1, 2026 snapshot and the Artificial Analysis Intelligence Index. Public figures are in the table below.

Source notes

Arena and Artificial Analysis use different weighting, so their placements won't necessarily match; the public figures are below.

SourceModel / tierPublic figures (approx.)How to read
Arena Text OverallClaude Fable 51509±6 Elo, No. 1 overallTop of the overall popularity leaderboard
Arena Text OverallClaude Opus 4.6 / 4.7 (thinking)Approx. 1505 / 1502closely behind; the same-generation non-thinking model also lands near the top ten
Arena Text OverallQwen3.8-max1496±10 (Preliminary)makes the overall top five; with a thin vote count and a wide interval, read it as an upward signal
Arena Text OverallGemini 3 / 3.1 Pro Preview, Kimi K3-max, GPT-5.6 Sol (xhigh)roughly 1483–1486the confidence intervals overlap heavily, so forcing a strict ranking isn't warranted
Arena Text OverallDeepSeek V4 Proaround 1458 Elostill a gap behind the top closed-source entries; pricing and self-hosting are separate considerations
AA Intelligence IndexClaude Opus 5 (max)≈61near the top of the composite index
AA Intelligence IndexOpus 5 (xhigh), Claude Fable 5Fable ≈60Close behind Opus max
AA Intelligence IndexGPT-5.6 Sol (max)≈59Right behind Anthropic's front-runners
AA Intelligence IndexKimi K3 (max) / DeepSeek V4 Pro (max)≈57 / ≈44Clear gaps within the open-source camp

July addendum: Kimi K3 has launched and gone open-source, and the GPT-5.6 family is rolling into the Arena one model at a time — with thin vote counts, the standings are unstable. For Chinese models we're not applying the old SuperCLUE May 2026 scores; instead, we're looking at the relative positions of Qwen / Kimi / DeepSeek / GLM / Doubao on the Arena.

Divided into S / A / B by scenario

The tiers are an editorial judgment, not officially certified:

  • S:In this scenario, the public leaderboards and our daily usage both hold up, so these can serve as the default go-to candidates.
  • A:Exceptionally strong in one area, or standing out on price-performance / open-source controllability, but with clear weaknesses (few votes / preliminary, expensive, slow, or missing modalities).
  • B:Gets the job done, suitable as a fallback or for specialized tasks; one screenshot isn't enough to elevate it to the company's sole default.

The four tables below are split by scenario. The "Basis" column only restates the Arena / AA signals and the editor's tiering rationale from above; it introduces no new benchmark scores.

Coding & Frontend Delivery

On Arena Code/WebDev Overall (Aug 1), Opus 5-max is around 1705 Elo, ranking first; Kimi K3-max is around 1676; Qwen3.8-max WebDev around 1668 (Preliminary); GPT-5.6 Sol around 1620. On the Text Coding sub-list, Fable 5 is around 1553±9still sits in the top tier. A strong WebDev ranking doesn't mean it's first at everything.

TierModelBasis / Leaderboard signalNotes
SClaude Opus 5 (max/high)WebDev Overall ≈1705, ranks firstDefault primary candidate; note that the max/high and chatbox tiers are not on the same row.
SClaude Fable 5Top five on WebDev; Text Coding ≈1553±9Dependable on both the coding-delivery and popularity sub-rankings.
Between S and AKimi K3-maxWebDev ≈1676; has beaten a host of closed-source flagships on the Fullstack sub-ranking.Prioritize A/B testing for frontend delivery and page-level redesigns; don't take this as an overall #1.
AGPT-5.6 Sol (xhigh, including Codex harness)WebDev ≈1620If the team already runs an OpenAI stack, do A/B testing; no need to immediately swap out the old default.
AQwen3.8-maxWebDev ≈1668 (Preliminary)Low vote count and limited reference value; teams on Tongyi can test it in parallel
BGLM-5.2-maxVisible in the top 15 for WebDevChina-market API compatibility and a cost fallback
BDeepSeek V4 Flash / Pro seriesOverall leaderboard standing and engineering costCheap and MIT-licensed for self-hosting; when closed-source flagships are still more stable, it serves as the high-traffic workhorse

Long-context writing, synthesis, and multimodal reading assistance

On crowd-pleasing dimensions such as Overall, Creative Writing, and Instruction Following, the Claude line has consistently ranked near the top; Gemini / GPT each have their own selling points in long context, multimodal, and product integration. Before picking a model, ask yourself: are you optimizing for blind-test perception or task pass rate?

TierModelBasis / leaderboard signalNotes
SClaude Fable 5 / Opus 4.x–5 familyRanks at the front in Overall and in popularity metrics for Writing and Instruction FollowingLong-document revision, edits to spec, fewer "greasy clichés" — the hands-on section is flagged as editorial opinion
AGemini 3 / 3.1 Pro PreviewText Overall around 1483–1486Long context and multimodal document streams are the usual selling points
AGemini 3.5 / 3.6 FlashSame Flash product lineA fast, cheap drafting machine that often makes more sense than the Pro
AGPT-5.5 / 5.6 familySol-tier AA index sits close to Anthropic's front row; Arena popularity rankings will be out of alignmentBroad tool ecosystem and product integration — just be clear about what you're optimizing for
BGrok 4.5, etc.Some agent/doc-oriented sub-leaderboards are more activeThe main Text overall leaderboard isn't necessarily #1 every day; move up a tier only if you have real-time retrieval or similar needs

Chinese-language business and open-source controllability

Don't blindly apply SuperCLUE's May scores. Check the relative positions of Qwen / Kimi / DeepSeek / GLM / Doubao on Arena, then layer on your business corpus. Being "second on the leaderboard" as an open-source model doesn't extrapolate to "runs fine in your server room."

TierModelBasis / Leaderboard signalsNotes
AQwen3.8-max / Qwen3.7 seriesArena Overall top five (few votes; limited reference value)A common default for Chinese product iteration; "Preliminary" shouldn't be taken as a final conclusion
AKimi K3Open-source weights + strong showing on coding leaderboardsCommon for Chinese-language business and controllable stacks; watch the MoE footprint and multi-node cost
ADeepSeek V4 Pro / FlashArena ≈1458; AA max ≈44MIT-licensed, low-priced, and self-hostable — what counts is throughput per dollar; closed-source S is not a free substitute
BGLM-5.x, MiMo, Doubao / Seed, etc.Can climb the Arena leaderboardDomestic API compatibility and a cost floor; before upgrading to S, run a pass against your own ticket/support/contract corpus

Tool calling and agentic direction

The AA Intelligence Index weighs agentic tasks more heavily; Computer Use / multi-step tool chains are Anthropic's product-side focus. If subscores haven't been independently rechecked, we won't invent percentages here. The most common failure: treating "a working demo" as "production-ready."

TierModelBasis / leaderboard signalsNotes
SClaude Opus 5 / Fable 5Near the top of the AA index (Opus 5-max ≈61; Fable ≈60)Default candidate for multi-step toolchains; no fabricated percentages for subscores that haven't been independently verified
AGPT-5.6 SolClosely tracks the index and harness entriesWhen you already have an OpenAI agent / Codex stack, prioritize side-by-side evaluation
AKimi K3Browsing, front-end delivery, and open-source controllabilityVery competitive when you need a self-hosted agent stack
BThe budget Flash / Luna / open-source small-tier optionsCost- and throughput-orientedSuits routing layers, pre-classification, and high-concurrency coarse processing; not suited as the sole production default

Five pitfalls to avoid when reading the leaderboard

  1. Style preferences don't extrapolate to correctness:The Arena is a blind human vote for which response people prefer. A model with a likable tone and polished formatting won't necessarily lead on rigorous reasoning tasks.
  2. Preliminary labels and wide confidence intervals:Models with few votes—some of the new max entries have only a few thousand—can jump three rungs in a day. When you see a ±15–20 interval, assume votes are still rolling in and the ranking isn't stable.
  3. Contamination and benchmark gaming:Public static benchmarks (MMLU-Pro, GPQA, etc.) get criticized for contamination year after year; we outright ignore any "perfect score" screenshots that lack a date and reproduction notes.
  4. A closed API isn't the same model as the chat version:Under the same brand, chat-latest, thinking, max effort, and codex-harness sit on different rows of the leaderboard. Write the wrong tier into a procurement contract and you've effectively bought a different model.
  5. The hidden costs of open-source weights:For an oversized MoE like K3, the weight footprint and multi-machine inference costs turn "second on the leaderboard" into "we can't run it in our data center." DeepSeek / GLM may actually be the more S-tier pick when it comes to deployability.

Modify our version with the template

The above is the editorial opinion of the OBPAI geek team. The template "AI Large Language Model (LLM) Overall Capability Ladder Chart" (ID 1785671363439) provides draggable tiles—just drag them to build the model leaderboard you think is right.

互动体验:拖拽排位《LLM Tier List -2024

LLM Tier List -2024

Unranked pool(0)
All items have been ranked!

For embedding steps, see How to Embed the Interactive Ladder Chart

How to Read the Public Leaderboard in Early August

In early August 2026, the overall popularity leaderboard is still crowded at the top with Anthropic models; on the Synthetic Intelligence Index, Opus 5, Fable 5, and GPT-5.6 Sol are bunched tightly together; and on the Coding Delivery leaderboard, both Kimi K3 and Opus 5 warrant serious A/B testing. When choosing, first define the use case — coding, long-form text, Chinese, or agent work — then decide whether you can live with closed-source models and the resulting bills. Rankings refresh weekly, so defer to the leaderboard you yourself pulled up.

榜单开放与生成

LLM Tier List -2024

想打造自己的个性化天梯图并开放给读者嵌入吗?

全屏高清制作
文章来源:OBPAI TierList