返回博客列表
Media Kit
Add an email on your account to be notified after we update
The OBPAI Geek Team•2026/9/24•0 阅读

Coding Model Rankings 2026

As of 2026-09-24. Terminal-Bench 4.0: GPT-6 Astra (max) with Codex resolves 58.2%. mini-SWE-agent 2.0.0 on SWE-bench Verified is a second table.

coding modelsTerminal-BenchSWE-benchGPT-62026
Coding Model Rankings 2026 - 天梯图与数据榜单

Compiled 2026-09-24. The main table copies the official Terminal-Bench 4.0 board; the board record was updated on 21 September 2026. Each row is one submission: model, reasoning effort and agent are scored together. GPT-6 Astra (max) with Codex resolves 58.2% of tasks; the 95% half-width is 2.8 points. Fable 5.1 (max) with Claude Code resolves 57.9%, half-width 3.8 points, and the two intervals overlap. Coding products: AI coding assistants. Crowd votes: LLM capability. List prices: LLM API pricing.

2026年全球编程大模型天梯

Terminal-Bench 4.0 official board

These are all 27 rows on that board. Tied ranks stay as printed. Cost is the bill for that evaluation run, not a per-million-token list price. GLM-5.3 (max) with Claude Code is 16th at 41.8%. Grok 4.7 (xhigh) with Grok Build was submitted on 21 September 2026 at 37.6%.

#模型档位代理机构解决率95% 半宽花费Token提交日
1GPT-6 AstramaxCodexOpenAI58.2%±2.8$3,2671.5BSep 3, 2026
2Fable 5.1maxClaude CodeAnthropic57.9%±3.8$6,2442.7BSep 1, 2026
2GPT-6 AstraxhighCodexOpenAI57.9%±2.7$2,3511.2BSep 3, 2026
2GPT-6 AstrahighCodexOpenAI57.9%±3.0$2,2691.2BSep 3, 2026
2Fable 5.1xhighClaude CodeAnthropic57.9%±3.4$4,8722.3BSep 1, 2026
6Fable 5.1highClaude CodeAnthropic54.5%±3.4$3,9852.2BSep 1, 2026
7GPT-6 AstramediumCodexOpenAI54.2%±2.7$1,9151.1BSep 3, 2026
8Fable 5.1mediumClaude CodeAnthropic53.9%±3.4$2,8331.6BSep 1, 2026
8Opus 5xhighClaude CodeAnthropic53.9%±3.2$6,0866.9BJul 24, 2026
10Opus 5maxClaude CodeAnthropic51.8%±3.4$5,9696.5BJul 24, 2026
11GPT-6 AstralowCodexOpenAI50.6%±2.8$1,557889.8MSep 3, 2026
12Opus 5highClaude CodeAnthropic50.3%±3.7$4,6625.5BJul 24, 2026
13Opus 5mediumClaude CodeAnthropic44.9%±3.8$3,1923.7BJul 24, 2026
14Fable 5maxClaude CodeAnthropic44.5%±3.9$7,2653.8BJun 9, 2026
15Fable 5.1lowClaude CodeAnthropic43.3%±3.6$2,3591.3BSep 1, 2026
16GLM-5.3maxClaude CodeZ.ai41.8%±3.2$2,7288.7BAug 14, 2026
17Grok 4.7xhighGrok BuildxAI37.6%±3.5$3,6835.5BSep 21, 2026
18GPT-5.6 SolmaxCodexOpenAI37.3%±3.8$2,5424.4BJun 26, 2026
19Opus 5lowClaude CodeAnthropic34.9%±3.9$2,3942.7BJul 24, 2026
20Opus 4.8maxClaude CodeAnthropic23.6%±3.6$6,4816.4BMay 28, 2026
21GPT-5.6 TerramaxCodexOpenAI21.5%±3.2$1,7345.7BJun 26, 2026
22Grok 4.6highGrok BuildxAI20.3%±3.1$3,5924.0BAug 12, 2026
23Gemini 3.8 Flashhighmini-SWE-agentGoogle19.1%±3.4$1,82917.2BSep 2, 2026
24GPT-5.6 LunamaxCodexOpenAI17.3%±2.9$34711.6BJun 26, 2026
25Grok 4.5highGrok BuildxAI12.4%±2.6$2,0943.4BJul 16, 2026
25Sonnet 5maxClaude CodeAnthropic12.4%±3.1$9,60421.6BJun 30, 2026
27Gemini 3.7 Flashhighmini-SWE-agentGoogle11.2%±2.5$1,26211.1BAug 13, 2026

SWE-bench Verified, one agent: mini-SWE-agent 2.0.0

The public SWE-bench board stacks different agents. The table below keeps bash-only submissions on mini-SWE-agent 2.0.0 and sorts them by resolve rate. Claude 4.5 Opus (high) is at 76.80%, average cost $0.75, submitted 17 February 2026. MiniMax M2.5 (high) and Gemini 3 Flash (high) are both at 75.80%. Kimi K2.5 (high) is at 70.80%, average $0.15. DeepSeek V3.2 (high) is at 70.00%.

#模型档位解决率平均花费提交日
1Claude 4.5 Opushigh76.80%$0.752026-02-17
2Gemini 3 Flashhigh75.80%$0.362026-02-17
3MiniMax M2.5high75.80%$0.072026-02-17
4Claude 4.6 Opus—75.60%$0.552026-02-17
5GLM 5high72.80%$0.532026-02-17
6GPT 5.2high72.80%$0.472026-02-17
7GPT 5.2 Codex—72.80%$0.452026-02-19
8Claude 4.5 Sonnethigh71.40%$0.662026-02-17
9Kimi K2.5high70.80%$0.152026-02-17
10DeepSeek V3.2high70.00%$0.452026-02-17
11Gemini 3 Prohigh69.60%$0.962026-02-26
12GPT 5 mini—56.20%$0.052026-02-17

LiveCodeBench default window (through 1 May 2025)

The default board lists 454 problems dated 1 August 2024 through 1 May 2025. o4-mini (high) has Pass@1 80.2, Easy 99.1, Hard 63.5. DeepSeek-R1-0528 is 5th at 73.1. That generation stops at the window’s end date. Astra, Fable 5.1 and Opus 5, released later, are absent.

#模型档位Pass@1EasyMediumHard
1o4-minihigh80.299.189.463.5
2o3high75.899.184.457.1
3o4-minimedium74.298.286.552.7
4Gemini 2.5 Pro06-0573.699.187.250.2
5DeepSeek-R1-0528—73.198.785.250.7
6Gemini 2.5 Pro05-0671.898.282.350.2
7EXAONE 4.0 32B—7098.482.346.2
8OpenReasoning-Nemotron-32B—69.898.381.446.3

How this list is ranked

Compiled 2026-09-24. The main table is the public board titled Terminal-Bench 4.0 on tbench.ai. The board record was updated on 21 September 2026. All 27 rows are included. Resolve rate, 95% half-width, cost, tokens and submission date are copied from that board’s JSON. Cost is total_cost_usd rounded to the dollar. The second table keeps only SWE-bench Verified rows that use bash-only and mini-SWE-agent 2.0.0, sorted by resolve rate. Those ranks are internal to the 12 rows. The third table is the LiveCodeBench default window, problems dated 1 August 2024 through 1 May 2025 (454 problems). That generation stops in May 2025 and is not folded into the main order.

SourceSnapshotWhat it measuresWeight
Terminal-Bench 4.0 official boardFetched 2026-09-24; board updated_at 21 Sep 2026Terminal-task resolve rate. One row is model + reasoning effort + agent. GPT-6 Astra (max) + Codex 58.2% (±2.8), 330 trials. All 27 rowsMain
SWE-bench Verified (bash-only)Fetched 2026-09-24; row dates 17–26 Feb 2026mini-SWE-agent 2.0.0 only. Claude 4.5 Opus (high) 76.80%, average $0.75. Other agents on the public board are left outAux
LiveCodeBenchDefault window 1 Aug 2024–1 May 2025; fetched 2026-09-24Contest-style Pass@1. o4-mini (high) 80.2. Models released after the window are absent from this default boardAux (older window)

互动体验:拖拽排位《Coding Models 2026》

榜单开放与生成

Coding Models 2026

想打造自己的个性化天梯图并开放给读者嵌入吗?

全屏高清制作
文章来源:OBPAI TierList