LLM API Pricing 2026
As of 2026-09-19. Official standard output: Qwen 3.7 Flash $0.130/M 1st, Ministral 3 14B $0.20 2nd, Qwen 3.8 Flash $0.47 3rd.

As of 19 September 2026. The main table uses each vendor’s official standard realtime list price in US dollars per million tokens, short context, cache miss, ranked by output price from low to high. Qwen 3.7 Flash is 1st at $0.130 per million output on the international sheet, Ministral 3 14B 2nd at $0.20, Qwen 3.8 Flash 3rd at $0.47. Capability: LLM rankings. Aggregators: model APIs and routers.

Official USD list prices
DeepSeek rows are off-peak. Gemini 3.8 Flash uses the rate in force through 31 December 2026. Gemini 3.1 Pro is marked Preview on the pricing page. When output prices match, the lower input price sits first.
| # | Model | Input $/M | Output $/M | Window |
|---|---|---|---|---|
| 1 | Qwen 3.7 Flash | $0.030 | $0.130 | Intl ≤32K |
| 2 | Ministral 3 14B | $0.2 | $0.2 | Standard |
| 3 | Qwen 3.8 Flash | $0.15 | $0.47 | Intl ≤1M |
| 4 | DeepSeek Flash | $0.15 | $0.60 | Off-peak, cache miss |
| 5 | Mistral Small 4 | $0.15 | $0.60 | Standard |
| 6 | GPT-5.6 Luna | $0.20 | $1.20 | Short-context standard |
| 7 | Mistral Large 3 | $0.5 | $1.5 | Standard |
| 8 | Qwen 3.7 Plus | $0.4 | $1.6 | Intl ≤256K |
| 9 | DeepSeek V4 Pro | $0.66 | $1.98 | Off-peak, cache miss |
| 10 | grok-build-0.1 | $1.00 | $2.00 | Short context |
| 11 | Gemini 3.5 Flash-Lite | $0.30 | $2.50 | Paid tier |
| 12 | grok-4.3 | $1.25 | $2.50 | Short context |
| 13 | Gemini 3.8 Flash | $0.75 | $3.75 | Current through 31 Dec 2026 |
| 14 | Claude Haiku 4.5 | $1 | $5 | Base input / output |
| 15 | grok-4.6 | $2.00 | $6.00 | Short context |
| 16 | Qwen 3.8 Max | $2 | $6 | Intl ≤1M |
| 17 | Gemini 3.5 Flash | $1.50 | $9.00 | Paid tier |
| 18 | Gemini 2.5 Pro | $1.25 | $10.00 | Prompts ≤200k |
| 19 | Claude Sonnet 5 | $2 | $10 | Base input / output |
| 20 | Gemini 3.1 Pro Preview | $2.00 | $12.00 | Preview; prompts ≤200k |
| 21 | GPT-5.6 Terra | $2.00 | $12.00 | Short-context standard |
| 22 | GPT-5.6 Sol | $4.00 | $20.00 | Short-context standard |
| 23 | Claude Opus 5 | $5 | $25 | Base input / output |
| 24 | Claude Fable 5.1 | $10 | $50 | Base input / output |
| 25 | GPT-6 Astra | $10.00 | $50.00 | Short-context standard |
How this list is ranked
As of 2026-09-19. The main table uses each vendor’s official list price for standard realtime inference in US dollars per million tokens, short context, cache miss. Ranks follow output price, low to high.
| Source | Snapshot | What it prints | Weight |
|---|---|---|---|
| Alibaba Cloud Model Studio pricing | 2026-09-19 | International: qwen3.7-flash ≤32K $0.030 / $0.130; qwen3.8-flash $0.15 / $0.47; qwen3.7-plus ≤256K $0.4 / $1.6; qwen3.8-max $2 / $6 | Main |
| Mistral Docs Pricing | 2026-09-19 | Standard: Ministral 3 14B $0.2 / $0.2; Small 4 $0.15 / $0.60; Large 3 $0.5 / $1.5 | Main |
| DeepSeek Models & Pricing | 2026-09-19 | Off-peak cache miss: Flash $0.15 / $0.60; V4 Pro $0.66 / $1.98. Peak is 2× | Main |
| OpenAI API Pricing | 2026-09-19 | Flagship short-context Standard: Luna $0.20 / $1.20; Terra $2 / $12; Sol $4 / $20; Astra $10 / $50. Not Batch / Flex / Fast | Main |
| xAI API Pricing | 2026-09-19 | Short context: grok-build-0.1 $1 / $2; grok-4.3 $1.25 / $2.50; grok-4.6 $2 / $6 | Main |
| Gemini Developer API pricing | 2026-09-19 | Paid standard: 3.5 Flash-Lite $0.30 / $2.50; 3.8 Flash current through 31 Dec 2026 $0.75 / $3.75; 3.5 Flash $1.50 / $9; 2.5 Pro ≤200k $1.25 / $10; 3.1 Pro Preview ≤200k $2 / $12 | Main |
| Claude Platform Docs: Pricing | 2026-09-19 | Base input / output: Haiku 4.5 $1 / $5; Sonnet 5 $2 / $10; Opus 5 $5 / $25; Fable 5.1 $10 / $50. Not Fast mode | Main |
| Zhipu API pricing | 2026-09-19 | CNY: GLM-5.3-Flash 0.8 / 2.8 yuan; GLM-5.3 8 / 28 yuan. No USD column | Aux (CNY) |
DeepSeek peak
The same page prints peak at 2× off-peak. Flash peak is $0.30 / $1.20; V4 Pro peak is $1.32 / $3.96. The main table uses off-peak only.
Zhipu yuan
Zhipu’s sheet is yuan per million tokens: GLM-5.3-Flash 0.8 input / 2.8 output; GLM-5.3 8 / 28. There is no USD column, so they stay off the dollar table.
| Model | Input | Output |
|---|---|---|
| GLM-5.3-Flash | 0.8 yuan / million | 2.8 yuan / million |
| GLM-5.3 | 8 yuan / million | 28 yuan / million |
互动体验:拖拽排位《LLM API Pricing 2026》
LLM API Pricing 2026
想打造自己的个性化天梯图并开放给读者嵌入吗?