返回博客列表
Add an email on your account to be notified after we update
The OBPAI Geek Team2026/9/190 阅读

LLM API Pricing 2026

As of 2026-09-19. Official standard output: Qwen 3.7 Flash $0.130/M 1st, Ministral 3 14B $0.20 2nd, Qwen 3.8 Flash $0.47 3rd.

LLMAPI pricingDeepSeekQwen2026
LLM API Pricing 2026 - 天梯图与数据榜单

As of 19 September 2026. The main table uses each vendor’s official standard realtime list price in US dollars per million tokens, short context, cache miss, ranked by output price from low to high. Qwen 3.7 Flash is 1st at $0.130 per million output on the international sheet, Ministral 3 14B 2nd at $0.20, Qwen 3.8 Flash 3rd at $0.47. Capability: LLM rankings. Aggregators: model APIs and routers.

2026年全球大模型API价格天梯

Official USD list prices

DeepSeek rows are off-peak. Gemini 3.8 Flash uses the rate in force through 31 December 2026. Gemini 3.1 Pro is marked Preview on the pricing page. When output prices match, the lower input price sits first.

#ModelInput $/MOutput $/MWindow
1Qwen 3.7 Flash$0.030$0.130Intl ≤32K
2Ministral 3 14B$0.2$0.2Standard
3Qwen 3.8 Flash$0.15$0.47Intl ≤1M
4DeepSeek Flash$0.15$0.60Off-peak, cache miss
5Mistral Small 4$0.15$0.60Standard
6GPT-5.6 Luna$0.20$1.20Short-context standard
7Mistral Large 3$0.5$1.5Standard
8Qwen 3.7 Plus$0.4$1.6Intl ≤256K
9DeepSeek V4 Pro$0.66$1.98Off-peak, cache miss
10grok-build-0.1$1.00$2.00Short context
11Gemini 3.5 Flash-Lite$0.30$2.50Paid tier
12grok-4.3$1.25$2.50Short context
13Gemini 3.8 Flash$0.75$3.75Current through 31 Dec 2026
14Claude Haiku 4.5$1$5Base input / output
15grok-4.6$2.00$6.00Short context
16Qwen 3.8 Max$2$6Intl ≤1M
17Gemini 3.5 Flash$1.50$9.00Paid tier
18Gemini 2.5 Pro$1.25$10.00Prompts ≤200k
19Claude Sonnet 5$2$10Base input / output
20Gemini 3.1 Pro Preview$2.00$12.00Preview; prompts ≤200k
21GPT-5.6 Terra$2.00$12.00Short-context standard
22GPT-5.6 Sol$4.00$20.00Short-context standard
23Claude Opus 5$5$25Base input / output
24Claude Fable 5.1$10$50Base input / output
25GPT-6 Astra$10.00$50.00Short-context standard

How this list is ranked

As of 2026-09-19. The main table uses each vendor’s official list price for standard realtime inference in US dollars per million tokens, short context, cache miss. Ranks follow output price, low to high.

SourceSnapshotWhat it printsWeight
Alibaba Cloud Model Studio pricing2026-09-19International: qwen3.7-flash ≤32K $0.030 / $0.130; qwen3.8-flash $0.15 / $0.47; qwen3.7-plus ≤256K $0.4 / $1.6; qwen3.8-max $2 / $6Main
Mistral Docs Pricing2026-09-19Standard: Ministral 3 14B $0.2 / $0.2; Small 4 $0.15 / $0.60; Large 3 $0.5 / $1.5Main
DeepSeek Models & Pricing2026-09-19Off-peak cache miss: Flash $0.15 / $0.60; V4 Pro $0.66 / $1.98. Peak is 2×Main
OpenAI API Pricing2026-09-19Flagship short-context Standard: Luna $0.20 / $1.20; Terra $2 / $12; Sol $4 / $20; Astra $10 / $50. Not Batch / Flex / FastMain
xAI API Pricing2026-09-19Short context: grok-build-0.1 $1 / $2; grok-4.3 $1.25 / $2.50; grok-4.6 $2 / $6Main
Gemini Developer API pricing2026-09-19Paid standard: 3.5 Flash-Lite $0.30 / $2.50; 3.8 Flash current through 31 Dec 2026 $0.75 / $3.75; 3.5 Flash $1.50 / $9; 2.5 Pro ≤200k $1.25 / $10; 3.1 Pro Preview ≤200k $2 / $12Main
Claude Platform Docs: Pricing2026-09-19Base input / output: Haiku 4.5 $1 / $5; Sonnet 5 $2 / $10; Opus 5 $5 / $25; Fable 5.1 $10 / $50. Not Fast modeMain
Zhipu API pricing2026-09-19CNY: GLM-5.3-Flash 0.8 / 2.8 yuan; GLM-5.3 8 / 28 yuan. No USD columnAux (CNY)

DeepSeek peak

The same page prints peak at 2× off-peak. Flash peak is $0.30 / $1.20; V4 Pro peak is $1.32 / $3.96. The main table uses off-peak only.

Zhipu yuan

Zhipu’s sheet is yuan per million tokens: GLM-5.3-Flash 0.8 input / 2.8 output; GLM-5.3 8 / 28. There is no USD column, so they stay off the dollar table.

ModelInputOutput
GLM-5.3-Flash0.8 yuan / million2.8 yuan / million
GLM-5.38 yuan / million28 yuan / million

互动体验:拖拽排位《LLM API Pricing 2026

榜单开放与生成

LLM API Pricing 2026

想打造自己的个性化天梯图并开放给读者嵌入吗?

全屏高清制作
文章来源:OBPAI TierList