2026 LLM Comprehensive Capability Leaderboard
Based on the Arena 2026-08-01 and Artificial Analysis public data from recent months, we assign S/A/B grades across coding, long-context, Chinese/open-source, and Agent scenarios, and include a draggable LLM leaderboard template for custom revisions.

How the ranking works
Data as of 2026-08-05Scenarios — coding, long-form writing, Chinese/open-source, and agent — are graded S/A/B, based mainly on the Arena (arena.ai) August 1, 2026 snapshot and the Artificial Analysis Intelligence Index. Public figures are in the table below.
Source notes
Arena and Artificial Analysis use different weighting, so their placements won't necessarily match; the public figures are below.
| Source | Model / tier | Public figures (approx.) | How to read |
|---|---|---|---|
| Arena Text Overall | Claude Fable 5 | 1509±6 Elo, No. 1 overall | Top of the overall popularity leaderboard |
| Arena Text Overall | Claude Opus 4.6 / 4.7 (thinking) | Approx. 1505 / 1502 | closely behind; the same-generation non-thinking model also lands near the top ten |
| Arena Text Overall | Qwen3.8-max | 1496±10 (Preliminary) | makes the overall top five; with a thin vote count and a wide interval, read it as an upward signal |
| Arena Text Overall | Gemini 3 / 3.1 Pro Preview, Kimi K3-max, GPT-5.6 Sol (xhigh) | roughly 1483–1486 | the confidence intervals overlap heavily, so forcing a strict ranking isn't warranted |
| Arena Text Overall | DeepSeek V4 Pro | around 1458 Elo | still a gap behind the top closed-source entries; pricing and self-hosting are separate considerations |
| AA Intelligence Index | Claude Opus 5 (max) | ≈61 | near the top of the composite index |
| AA Intelligence Index | Opus 5 (xhigh), Claude Fable 5 | Fable ≈60 | Close behind Opus max |
| AA Intelligence Index | GPT-5.6 Sol (max) | ≈59 | Right behind Anthropic's front-runners |
| AA Intelligence Index | Kimi K3 (max) / DeepSeek V4 Pro (max) | ≈57 / ≈44 | Clear gaps within the open-source camp |
July addendum: Kimi K3 has launched and gone open-source, and the GPT-5.6 family is rolling into the Arena one model at a time — with thin vote counts, the standings are unstable. For Chinese models we're not applying the old SuperCLUE May 2026 scores; instead, we're looking at the relative positions of Qwen / Kimi / DeepSeek / GLM / Doubao on the Arena.
Divided into S / A / B by scenario
The tiers are an editorial judgment, not officially certified:
- S:In this scenario, the public leaderboards and our daily usage both hold up, so these can serve as the default go-to candidates.
- A:Exceptionally strong in one area, or standing out on price-performance / open-source controllability, but with clear weaknesses (few votes / preliminary, expensive, slow, or missing modalities).
- B:Gets the job done, suitable as a fallback or for specialized tasks; one screenshot isn't enough to elevate it to the company's sole default.
The four tables below are split by scenario. The "Basis" column only restates the Arena / AA signals and the editor's tiering rationale from above; it introduces no new benchmark scores.
Coding & Frontend Delivery
On Arena Code/WebDev Overall (Aug 1), Opus 5-max is around 1705 Elo, ranking first; Kimi K3-max is around 1676; Qwen3.8-max WebDev around 1668 (Preliminary); GPT-5.6 Sol around 1620. On the Text Coding sub-list, Fable 5 is around 1553±9still sits in the top tier. A strong WebDev ranking doesn't mean it's first at everything.
| Tier | Model | Basis / Leaderboard signal | Notes |
|---|---|---|---|
| S | Claude Opus 5 (max/high) | WebDev Overall ≈1705, ranks first | Default primary candidate; note that the max/high and chatbox tiers are not on the same row. |
| S | Claude Fable 5 | Top five on WebDev; Text Coding ≈1553±9 | Dependable on both the coding-delivery and popularity sub-rankings. |
| Between S and A | Kimi K3-max | WebDev ≈1676; has beaten a host of closed-source flagships on the Fullstack sub-ranking. | Prioritize A/B testing for frontend delivery and page-level redesigns; don't take this as an overall #1. |
| A | GPT-5.6 Sol (xhigh, including Codex harness) | WebDev ≈1620 | If the team already runs an OpenAI stack, do A/B testing; no need to immediately swap out the old default. |
| A | Qwen3.8-max | WebDev ≈1668 (Preliminary) | Low vote count and limited reference value; teams on Tongyi can test it in parallel |
| B | GLM-5.2-max | Visible in the top 15 for WebDev | China-market API compatibility and a cost fallback |
| B | DeepSeek V4 Flash / Pro series | Overall leaderboard standing and engineering cost | Cheap and MIT-licensed for self-hosting; when closed-source flagships are still more stable, it serves as the high-traffic workhorse |
Long-context writing, synthesis, and multimodal reading assistance
On crowd-pleasing dimensions such as Overall, Creative Writing, and Instruction Following, the Claude line has consistently ranked near the top; Gemini / GPT each have their own selling points in long context, multimodal, and product integration. Before picking a model, ask yourself: are you optimizing for blind-test perception or task pass rate?
| Tier | Model | Basis / leaderboard signal | Notes |
|---|---|---|---|
| S | Claude Fable 5 / Opus 4.x–5 family | Ranks at the front in Overall and in popularity metrics for Writing and Instruction Following | Long-document revision, edits to spec, fewer "greasy clichés" — the hands-on section is flagged as editorial opinion |
| A | Gemini 3 / 3.1 Pro Preview | Text Overall around 1483–1486 | Long context and multimodal document streams are the usual selling points |
| A | Gemini 3.5 / 3.6 Flash | Same Flash product line | A fast, cheap drafting machine that often makes more sense than the Pro |
| A | GPT-5.5 / 5.6 family | Sol-tier AA index sits close to Anthropic's front row; Arena popularity rankings will be out of alignment | Broad tool ecosystem and product integration — just be clear about what you're optimizing for |
| B | Grok 4.5, etc. | Some agent/doc-oriented sub-leaderboards are more active | The main Text overall leaderboard isn't necessarily #1 every day; move up a tier only if you have real-time retrieval or similar needs |
Chinese-language business and open-source controllability
Don't blindly apply SuperCLUE's May scores. Check the relative positions of Qwen / Kimi / DeepSeek / GLM / Doubao on Arena, then layer on your business corpus. Being "second on the leaderboard" as an open-source model doesn't extrapolate to "runs fine in your server room."
| Tier | Model | Basis / Leaderboard signals | Notes |
|---|---|---|---|
| A | Qwen3.8-max / Qwen3.7 series | Arena Overall top five (few votes; limited reference value) | A common default for Chinese product iteration; "Preliminary" shouldn't be taken as a final conclusion |
| A | Kimi K3 | Open-source weights + strong showing on coding leaderboards | Common for Chinese-language business and controllable stacks; watch the MoE footprint and multi-node cost |
| A | DeepSeek V4 Pro / Flash | Arena ≈1458; AA max ≈44 | MIT-licensed, low-priced, and self-hostable — what counts is throughput per dollar; closed-source S is not a free substitute |
| B | GLM-5.x, MiMo, Doubao / Seed, etc. | Can climb the Arena leaderboard | Domestic API compatibility and a cost floor; before upgrading to S, run a pass against your own ticket/support/contract corpus |
Tool calling and agentic direction
The AA Intelligence Index weighs agentic tasks more heavily; Computer Use / multi-step tool chains are Anthropic's product-side focus. If subscores haven't been independently rechecked, we won't invent percentages here. The most common failure: treating "a working demo" as "production-ready."
| Tier | Model | Basis / leaderboard signals | Notes |
|---|---|---|---|
| S | Claude Opus 5 / Fable 5 | Near the top of the AA index (Opus 5-max ≈61; Fable ≈60) | Default candidate for multi-step toolchains; no fabricated percentages for subscores that haven't been independently verified |
| A | GPT-5.6 Sol | Closely tracks the index and harness entries | When you already have an OpenAI agent / Codex stack, prioritize side-by-side evaluation |
| A | Kimi K3 | Browsing, front-end delivery, and open-source controllability | Very competitive when you need a self-hosted agent stack |
| B | The budget Flash / Luna / open-source small-tier options | Cost- and throughput-oriented | Suits routing layers, pre-classification, and high-concurrency coarse processing; not suited as the sole production default |
Five pitfalls to avoid when reading the leaderboard
- Style preferences don't extrapolate to correctness:The Arena is a blind human vote for which response people prefer. A model with a likable tone and polished formatting won't necessarily lead on rigorous reasoning tasks.
- Preliminary labels and wide confidence intervals:Models with few votes—some of the new max entries have only a few thousand—can jump three rungs in a day. When you see a ±15–20 interval, assume votes are still rolling in and the ranking isn't stable.
- Contamination and benchmark gaming:Public static benchmarks (MMLU-Pro, GPQA, etc.) get criticized for contamination year after year; we outright ignore any "perfect score" screenshots that lack a date and reproduction notes.
- A closed API isn't the same model as the chat version:Under the same brand, chat-latest, thinking, max effort, and codex-harness sit on different rows of the leaderboard. Write the wrong tier into a procurement contract and you've effectively bought a different model.
- The hidden costs of open-source weights:For an oversized MoE like K3, the weight footprint and multi-machine inference costs turn "second on the leaderboard" into "we can't run it in our data center." DeepSeek / GLM may actually be the more S-tier pick when it comes to deployability.
Modify our version with the template
The above is the editorial opinion of the OBPAI geek team. The template "AI Large Language Model (LLM) Overall Capability Ladder Chart" (ID 1785671363439) provides draggable tiles—just drag them to build the model leaderboard you think is right.
互动体验:拖拽排位《LLM Tier List -2024》
LLM Tier List -2024
For embedding steps, see How to Embed the Interactive Ladder Chart。
How to Read the Public Leaderboard in Early August
In early August 2026, the overall popularity leaderboard is still crowded at the top with Anthropic models; on the Synthetic Intelligence Index, Opus 5, Fable 5, and GPT-5.6 Sol are bunched tightly together; and on the Coding Delivery leaderboard, both Kimi K3 and Opus 5 warrant serious A/B testing. When choosing, first define the use case — coding, long-form text, Chinese, or agent work — then decide whether you can live with closed-source models and the resulting bills. Rankings refresh weekly, so defer to the leaderboard you yourself pulled up.
LLM Tier List -2024
想打造自己的个性化天梯图并开放给读者嵌入吗?