Large Model API Pricing

Explore the pricing for our model API. With transparent rates and flexible options, find the right plan to meet your needs.

Anthropic logo

Anthropic

Anthropic's Claude model offers advanced AI safety capabilities, focusing on useful, harmless, and honest AI assistants with powerful reasoning and conversational abilities.

Model NameInput Token RangeContextInput (/Mt)Cache Write (/Mt)Cache read (/Mt)Output (/Mt)Operation
claude-fable-5-1Official resource-1M$10$12.5(5m)·$20(1h)$0.25$50
claude-fable-5Official resource-1M$10$12.5(5m)·$20(1h)$1$50
claude-opus-4-7Official resource-1M
$4.75 $5
$5.9375 (5 min) · $9.50 (1 hr) $6.25 (5 min) × $10 (1 hr)
$0.475 $0.5
$23.75 $25
claude-opus-5Official resource-1M$5$6.25 (5 min) · $10 (1 hr)$0.5$25
claude-sonnet-5Official resource-1M$2$2.5(5m)·$4(1h)$0.2$10
claude-opus-4-8Official resource-1M
$4.75 $5
$5.9375 (5 min) · $9.50 (1 hr) $6.25 (5 min) × $10 (1 hr)
$0.475 $0.5
$23.75 $25
claude-opus-4-8-rValue zone-1M
$1$5
$1.25(5m)·$2(1h)$6.25(5m)·$10(1h)
$0.1$0.5
$5$25
claude-opus-4-7-rValue zone-1M
$1$5
$1.25(5m)·$2(1h)$6.25(5m)·$10(1h)
$0.1$0.5
$5$25
claude-opus-4-6-ddValue zone-1M
$2.75$5
$3.4375(5m)·$5.5(1h)$6.25(5m)·$10(1h)
$0.275$0.5
$13.75$25
claude-opus-4-6Official resource1-200K1M$5$6.25 (5 min) · $10 (1 hr)$0.5$25
200K-1M1M$5$6.25 (5 min) · $10 (1 hr)$0.5$25
claude-opus-4-6-rValue zone-1M
$1$5
$1.25(5m)·$2(1h)$6.25(5m)·$10(1h)
$0.1$0.5
$5$25
claude-sonnet-4-6Official resource1-200K1M$3$3.75 (5 min) · $6 (1 hr)$0.3$15
200K-1M1M$3$3.75 (5 min) · $6 (1 hr)$0.3$15
claude-sonnet-4-6-ddValue zone-1M
$1.65$3
$2.0625(5m)·$3.3(1h)$3.75(5m)·$6(1h)
$0.165$0.3
$8.25$15
claude-sonnet-4-6-rValue zone-1M
$0.6$3
$0.75(5m)·$1.2(1h)$3.75(5m)·$6(1h)
$0.06$0.3
$3$15
claude-opus-4-5-20251101Official resource-200K
$4.75 $5
$5.9375 (5 min) · $9.50 (1 hr) $6.25 (5 min) × $10 (1 hr)
$0.475 $0.5
$23.75 $25
claude-opus-4-5-20251101-ddValue zone-200K
$2.75$5
$3.4375(5m)$6.25(5m)
$0.275$0.5
$13.75$25
claude-sonnet-4-5-20250929Official resource1-200K200K$3$3.75 (5 min) · $6 (1 hr)$0.3$15
200K-1M200K$6$7.5 (5 min) × $12 (1 hr)$0.6$22.5
claude-sonnet-4-5-20250929-ddValue zone-200K
$1.65$3
$2.0625(5m)$3.75(5m)
$0.165$0.3
$8.25$15
claude-haiku-4-5-20251001Official resource-20K$1$1.25 (5 min) · $2 (1 hr)$0.1$5
claude-haiku-4-5-20251001-ddValue zone-200K
$0.55$1
$0.6875(5m)·$1.1(1h)$1.25(5m)·$2(1h)
$0.055$0.1
$2.75$5
claude-haiku-4-5-20251001-rValue zone-200K
$0.2$1
$0.25(5m)·$0.4(1h)$1.25(5m)·$2(1h)
$0.02$0.1
$1$5
claude-opus-4-8-ccValue zone-1M
$1.85$5
$2.3125(5m)·$3.7(1h)$6.25(5m)·$10(1h)
$0.185$0.5
$9.25$25
claude-opus-4-6-ccValue zone-1M
$1.85$5
$2.3125(5m)·$3.7(1h)$6.25(5m)·$10(1h)
$0.185$0.5
$9.25$25
claude-opus-4-7-ccValue zone-1M
$1.85$5
$2.3125(5m)·$3.7(1h)$6.25(5m)·$10(1h)
$0.185$0.5
$9.25$25
claude-sonnet-4-6-ccValue zone-1M
$1.11$3
$1.3875(5m)·$2.22(1h)$3.75(5m)·$6(1h)
$0.111$0.3
$5.55$15
claude-haiku-4-5-20251001-ccValue zone-200K
$0.37$1
$0.4625(5m)·$0.74(1h)$1.25(5m)·$2(1h)
$0.037$0.1
$1.85$5
OpenAI

OpenAI

OpenAI's GPT series of models offer state-of-the-art language understanding and generation capabilities, delivering outstanding performance across a wide range of tasks, and are among the industry's leading AI models.

Model NameInput Token RangeContextInput (/Mt)Cache Write (/Mt)Cache read (/Mt)Output (/Mt)Operation
gpt-6-astraOfficial resource1-272K1.1M
$9.5 $10
$11.875(30m)$12.5(30m)
$0.95$1
$47.5$50
272K-1.1M1.1M
$19$20
$23.75(30m)$25(30m)
$1.9 $2
$71.25 $75
gpt-5.6-terra-esValue zone1-272K3.7M
$0.2$2
$0.25(30m)$2.5(30m)
$0.02$0.2
$1.2$12
272K-372K3.7M
$0.4$4
$0.5(30m)$5(30m)
$0.04$0.4
$1.8$18
gpt-5.6-luna-esValue zone1-272K372K
$0.02$0.2
$0.025(30m)$0.25(30m)
$0.002$0.02
$0.12$1.2
272K-372K372K
$0.04$0.4
$0.05(30m)$0.5(30m)
$0.004$0.04
$0.18$1.8
gpt-5.6-sol-esValue zone1-272K372K
$0.4$4
$0.5(30m)$5(30m)
$0.04$0.4
$2$20
272K-372K372K
$0.8$8
$1(30m)$10(30m)
$0.08$0.8
$3$30
gpt-5.6-terraOfficial resource1-272K1.1M$2$2.5(30m)$0.2$12
272K-1.1M1.1M$4$5(30m)$0.4$18
gpt-5.6-lunaOfficial resource1-272K1.1M$0.2$0.25(30m)$0.02$1.2
272K-1.1M1.1M$0.4$0.5(30m)$0.04$1.8
gpt-5.6-solOfficial resource1-272K1.1M$4$5(30m)$0.4$20
272K-1.1M1.1M$8$10(30m)$0.8$30
gpt-5.5-rValue zone-1.1M
$0.5$5
-
$0.05$0.5
$3$30
gpt-5.5Official resource1-272K1.1M
$4.75 $5
-
$0.475 $0.5
$28.5$30
272K-1.1M1.1M
$9.5 $10
-
$0.95$1
$42.75$45
gpt-4.1-nanoOfficial resource-1M
$0.095 $0.1
-
$0.0237 $0.025
$0.38 $0.4
gpt-5.5-proOfficial resource1-272K1.1M$30--$180
272K-1.1M1.1M$60--$270
gpt-5.4-proOfficial resource1-272K1.1M$30--$180
272K-1.1M1.1M$60--$270
gpt-5.4Official resource1-272K1.1M$2.5-$0.25$15
272K-1.1M1.1M$5-$0.5$22.5
gpt-5.4-miniOfficial resource-400K
$0.7125 $0.75
-
$0.0712 $0.075
$4.275 $4.5
gpt-5.4-nanoOfficial resource-400K
$0.19 $0.2
-
$0.019 $0.02
$1.1875 $1.25
gpt-5.3-codexOfficial resource-400K
$1.6625 $1.75
-
$0.1662 $0.175
$13.3 $14
gpt-5.2-proOfficial resource-400K
$19.95 $21
--
$159.6 $168
gpt-5.2-codexOfficial resource-400K$1.75-$0.175$14
gpt-5.2Official resource-400K
$1.6625 $1.75
-
$0.1662 $0.175
$13.3 $14
gpt-5.1-codex-maxOfficial resource-400K
$1.1875 $1.25
-
$0.1187 $0.125
$9.5 $10
gpt-5.1-codexOfficial resource-400K
$1.1875 $1.25
-
$0.1187 $0.125
$9.5 $10
gpt-5.1-codex-miniOfficial resource-400K
$0.2375 $0.25
-
$0.0237 $0.025
$1.9 $2
gpt-5.1Official resource-400K
$1.1875 $1.25
-
$0.1187 $0.125
$9.5 $10
gpt-5-proOfficial resource-400K
$14.25 $15
--
$114 $120
gpt-5-codexOfficial resource-400K
$1.1875 $1.25
-
$0.1187 $0.125
$9.5 $10
gpt-5Official resource-400K
$1.1875 $1.25
-
$0.1187 $0.125
$9.5 $10
gpt-5-miniOfficial resource-400K
$0.2375 $0.25
-
$0.0237 $0.025
$1.9 $2
gpt-5-nanoOfficial resource-400K
$0.0475 $0.05
-
$0.0047 $0.005
$0.38 $0.4
gpt-4.1Official resource-1M$2-$0.5$8
gpt-4.1-miniOfficial resource-1M$0.4-$0.1$1.6
gpt-4oOfficial resource-131.1K
$2.375 $2.5
-
$1.1875 $1.25
$9.5 $10
gpt-4o-miniOfficial resource-128K
$0.1425 $0.15
-
$0.0712 $0.075
$0.57 $0.6
OpenAI: GPT OSS 20B-131.1K$0.05--$0.2
OpenAI GPT OSS 120B-131.1K$0.1--$0.5
Gemini logo

Gemini

Google's Gemini model offers high-quality natural language processing capabilities, performs exceptionally well across a wide range of NLP tasks, and boasts powerful multimodal capabilities.

Model NameInput Token RangeContextInput (/Mt)Cache Write (/Mt)Cache read (/Mt)Output (/Mt)Operation
gemini-3.1-flash-lite-1M$0.25$0.083(5m)·$1(1h)$0.025$1.5
gemini-3.5-flashOfficial resource-1M
$1.425 $1.5
$0.0788 (5 min) × $0.95 (1 hr) $0.083 (5 min) × $1 (1 hr)
$0.1425 $0.15
$8.55$9
gemini-3.1-pro-previewOfficial resource1-204.8K1M$2$0.375 (5 min) × $4.5 (1 hr)$0.2$12
204.8K-1M1M$4$0.375 (5 min) × $4.5 (1 hr)$0.4$18
gemini-3-flash-previewOfficial resource-1M
$0.475 $0.5
$0.0788 (5 min) × $0.95 (1 hr) $0.083 (5 min) × $1 (1 hr)
$0.0475 $0.05
$2.85 $3
gemini-2.5-proOfficial resource-1M
$1.1875 $1.25
$0.3562 (5-month) × $4.275 (1-hour) $0.375 (5 min) × $4.5 (1 hr)
$0.1187 $0.125
$9.5 $10
gemini-2.5-pro-preview-06-05Official resource-1M
$1.1875 $1.25
$0.3562 (5-month) × $4.275 (1-hour) $0.375 (5 min) × $4.5 (1 hr)
$0.1187 $0.125
$9.5 $10
gemini-2.5-flash-preview-05-20Official resource-1M
$0.1425 $0.15
$0.0788 (5 min) × $0.95 (1 hr) $0.083 (5 min) × $1 (1 hr)
$0.0285 $0.03
$3.325 $3.5
gemini-2.5-flash-1M
$0.285 $0.3
$0.0788 (5 min) × $0.95 (1 hr) $0.083 (5 min) × $1 (1 hr)
$0.0285 $0.03
$2.375 $2.5
gemini-2.5-flash-lite-preview-09-2025Official resource-1M
$0.095 $0.1
$0.0788 (5 min) × $0.95 (1 hr) $0.083 (5 min) × $1 (1 hr)
$0.0095 $0.01
$0.38 $0.4
gemini-2.5-flash-liteOfficial resource-1M
$0.095 $0.1
$0.0788 (5 min) × $0.95 (1 hr) $0.083 (5 min) × $1 (1 hr)
$0.0095 $0.01
$0.38 $0.4
Gemma 3 27B-32.8K$0.119--$0.2
Gemma3 12B-131.1K$0.05--$0.1
gemini-3.8-flash-1M$0.75$0.0417(5m)·$0.5(1h)$0.075$3.75
gemini-3.7-flash-1M$0.75$0.0417(5m)·$0.5(1h)$0.075$3.75
gemini-3.6-flashOfficial resource-1M$0.75$0.0417(5m)·$0.5(1h)$0.075$3.75
gemini-3.5-flash-liteOfficial resource-1M$0.3$0.0833(5m)·$1(1h)$0.03$2.5
Llama logo

Llama

Meta's Llama model offers state-of-the-art language understanding capabilities and features an open architecture, making it suitable for a wide range of applications.

Model NameContextInput (/Mt)Output (/Mt)Operation
Llama 3.1 8B Instruct16.4K$0.02$0.05
Llama 3.3 70B Instruct131.1K$0.13$0.39
Llama 4 Maverick Instructions1M$0.17$0.85
Llama 4 Scout Instructor131.1K$0.1$0.5
Qwen logo

Qwen

The Qwen series of models offers powerful natural language processing capabilities and is available in a range of parameter sizes, from lightweight to enterprise-grade solutions.

Model NameInput Token RangeContextInput (/Mt)Cache read (/Mt)Output (/Mt)Operation
Qwen3.5-Plus1-256K1M$0.4-$2.4
256K-1M1M$1.2-$7.2
Qwen3.8 Flash-1M$0.15$0.016$0.47
Qwen3.8 27B-1M$0.42$0.085$3
Qwen3.8 2.4T A95B-1M$2$0.25$6
Qwen3 235B A22B Instruct 2507-131.1K$0.15-$0.8
Qwen 2.5 72B Instruct-32K$0.38-$0.4
Qwen MT Plus-4.1K$0.25-$0.75
Qwen 2.5 7B Instruct-32K$0.07-$0.07
Qwen 2.5 VL 72B Instruction Manual-32.8K$0.8-$0.8
Qwen3 30B A3B-41K$0.09-$0.45
Qwen3 32B-41K$0.1-$0.45
Qwen3 235B A22B-41K$0.2-$0.8
Qwen3 235B A22b Thinking 2507-131.1K$0.3-$3
Qwen3 Coder 480B A35B Instructions-262.1K$0.38-$1.55
Qwen3 Coder Next FP8-262.1K$0.2-$1.5
Qwen3 Next 80B A3B Instruct-65.5K$0.15-$1.5
Qwen3 Next 80B A3B Thinking-65.5K$0.15-$1.5
Qwen3.5-27B-262.1K$0.3-$2.4
Qwen3.8 Max-1M$2$0.25$6
Qwen3.5-122B-A10B-262.1K$0.4-$3.2
Qwen3.5-35B-A3B-262.1K$0.25-$2
Qwen3.5-397B-A17B-262.1K$0.6-$3.6
Wenxin

Baidu

Baidu's ERNIE model offers advanced Chinese language understanding and multimodal capabilities, is optimized for Chinese applications, and is competitively priced.

Model NameContextInput (/Mt)Output (/Mt)Operation
ERNIE 4.5 VL 424B A47B123K$0.42$1.25
ERNIE 4.5 300B A47B123K$0.28$1.1
ChatGLM

THUDM

The GLM series of models from Tsinghua University feature advanced Chinese language understanding and generation capabilities.

Model NameContextInput (/Mt)Cache read (/Mt)Output (/Mt)Operation
GLM 5.3 Flash1M$0.15$0.03$0.5
GLM 5.31M$1.4$0.26$4.4
GLM 5.21M$1.4$0.26$4.4
GLM-5.1204.8K$1.38$0.26$4.4
GLM-5V-Turbo204.8K$1.2$0.24$4
GLM 4.5V65.5K$0.6-$1.8
GLM-4.5131.1K$0.6-$2.2
GLM-4.7204.8K$0.6-$2.2
GLM-4.7-Flash200K$0.07$0.01$0.4
GLM-5204.8K$1$0.2$3.2
GLM-5-Turbo202.8K$1.2$0.24$4
Sao10K logo

Sao10K

A fine-tuned model specifically optimized for creative and role-playing applications, featuring enhanced storytelling capabilities.

Model NameContextInput (/Mt)Output (/Mt)Operation
L3 8B Stheno V3.28.2K$0.05$0.05
L3 70B Euryale V2.1 8.2K$1.48$1.48
L31 70B Euryale V2.28.2K$1.48$1.48
Sao10k L3 8B Lunaris 8.2K$0.05$0.05
Mistralai logo

Mistralai

A powerful and efficient language model from Mistral AI, designed for both commercial and open-source applications.

Model NameContextInput (/Mt)Output (/Mt)Operation
Mistral Nemo60.3K$0.04$0.17
Mistral 7B Instruct32.8K$0.029$0.059
Deepseek logo

Deepseek

Advanced AI models from DeepSeek, offering cutting-edge inference capabilities and competitive pricing for enterprise and research applications.

Model NameContextInput (/Mt)Cache Write (/Mt)Cache read (/Mt)Output (/Mt)Operation
DeepSeek V4.1 Flash1M$0.3-$0.006$1.2
DeepSeek V4 Pro 08131M$1.32-$0.044$3.96
DeepSeek V4 Flash Vision Exp1M$0.44-$0.028$1.32
Deepseek V4 Flash 07311M$0.44-$0.028$1.32
Deepseek V4 Flash1M$0.14-$0.028$0.28
Deepseek V4 Pro1M$1.6-$0.135$3.2
DeepSeek R1 0528163.8K$0.7-$0.35$2.5
DeepSeek V3 0324163.8K$0.28$0.14 (5m)$0.14$1.14
DeepSeek V3.1163.8K$0.27--$1
DeepSeek-OCR 28.2K$0.03--$0.03
MiniMax logo

MiniMax

MiniMax AI's advanced language model delivers powerful conversational AI capabilities, excelling in customer service, content generation, and creative applications, with robust multilingual support and enterprise-grade scalability.

Model NameContextInput (/Mt)Output (/Mt)Operation
MiniMax M11M$0.55$2.2
Gryphe logo

Gryphe

An innovative AI model from Gryphe that offers professional-grade language understanding capabilities, with a focus on efficiency and adaptability, making it ideal for niche applications.

Model NameContextInput (/Mt)Output (/Mt)Operation
Mythomax L2 13B4.1K$0.09$0.09

Mixture of Experts

A sophisticated collection of state-of-the-art AI models, featuring advanced reasoning and mathematical proof capabilities, as well as cutting-edge language understanding across multiple domains.

Model NameInput Token RangeContextInput (/Mt)Cache Write (/Mt)Cache read (/Mt)Output (/Mt)Operation
Qwen3.5-Plus1-256K1M$0.4--$2.4
256K-1M1M$1.2--$7.2
DeepSeek V4.1 Flash-1M$0.3-$0.006$1.2
GLM 5.3 Flash-1M$0.15-$0.03$0.5
GLM 5.3-1M$1.4-$0.26$4.4
DeepSeek V4 Pro 0813-1M$1.32-$0.044$3.96
Qwen3.8 Flash-1M$0.15-$0.016$0.47
DeepSeek V4 Flash Vision Exp-1M$0.44-$0.028$1.32
Deepseek V4 Flash 0731-1M$0.44-$0.028$1.32
Qwen3.8 2.4T A95B-1M$2-$0.25$6
GLM 5.2-1M$1.4-$0.26$4.4
Kimi K2.7 Code-262.1K$0.95-$0.19$4
MiniMax M31-524.3K1M$0.3-$0.06$1.2
524.3K-1M1M$1.2-$0.24$4.8
Deepseek V4 Flash-1M$0.14-$0.028$0.28
Deepseek V4 Pro-1M$1.6-$0.135$3.2
MiniMax M2.7-204.8K$0.3-$0.03$1.2
Kimi K2.5-262.1K$0.6-$0.1$3
Kimi K2 Instruct-131.1K$0.57--$2.3
GLM-5.1-204.8K$1.38-$0.26$4.4
GLM-5V-Turbo-204.8K$1.2-$0.24$4
ERNIE 4.5 VL 424B A47B-123K$0.42--$1.25
OpenAI: GPT OSS 20B-131.1K$0.05--$0.2
OpenAI GPT OSS 120B-131.1K$0.1--$0.5
DeepSeek R1 0528-163.8K$0.7-$0.35$2.5
DeepSeek V3 0324-163.8K$0.28$0.14 (5m)$0.14$1.14
DeepSeek V3.1-163.8K$0.27--$1
ERNIE 4.5 300B A47B-123K$0.28--$1.1
GLM 4.5V-65.5K$0.6--$1.8
GLM-4.5-131.1K$0.6--$2.2
Kimi K2.6-262.1K$0.8-$0.16$3.4
GLM-4.7-204.8K$0.6--$2.2
GLM-4.7-Flash-200K$0.07-$0.01$0.4
GLM-5-204.8K$1-$0.2$3.2
GLM-5-Turbo-202.8K$1.2-$0.24$4
Llama 4 Maverick Instructions-1M$0.17--$0.85
Llama 4 Scout Instructor-131.1K$0.1--$0.5
MiniMax M1-1M$0.55--$2.2
Minimax M2.1-204.8K$0.3$0.375 (5m)$0.03$1.2
MiniMax M2.5-204.8K$0.3-$0.03$1.2
MiniMax M2.7-highspeed-204.8K$0.6-$0.06$2.4
MiniMax M2.5-highspeed-204.8K$0.6-$0.03$2.4
Qwen3 30B A3B-41K$0.09--$0.45
Qwen3 32B-41K$0.1--$0.45
Qwen3 235B A22B-41K$0.2--$0.8
Qwen3 235B A22b Thinking 2507-131.1K$0.3--$3
XiaomiMiMo/MiMo-V2-Pro1-262.1K1M$1-$0.2$3
262.1K-1M1M$2-$0.4$6
Qwen3.8 Max-1M$2-$0.25$6
XiaomiMiMo/MiMo-V2.5-Pro1-262.1K1M$1-$0.2$3
262.1K-1M1M$2-$0.4$6
Qwen3.5-122B-A10B-262.1K$0.4--$3.2
Qwen3.5-35B-A3B-262.1K$0.25--$2
Qwen3.5-397B-A17B-262.1K$0.6--$3.6
Contact Us