# CostOfToken — complete LLM pricing table

> Compare LLM API pricing across OpenAI, Anthropic, Google, xAI, DeepSeek, Qwen, Kimi and more. Input, cached and output cost per 1M tokens, normalized to USD and updated daily. Free public JSON API, no signup.

Last updated: 2026-08-11T18:01:39.601Z
Models: 216
Units: USD per 1,000,000 tokens, standard tier.

A price of `Free` means genuinely zero, not unknown; unknown is shown as `—`.
Batch, Flex and Priority tiers are excluded because they are not comparable to
other vendors' standard rates. Rows marked "via OpenRouter" come from a
reseller catalogue rather than the vendor's own page and may differ.

Machine-readable equivalent: https://www.costoftoken.com/api/v1/prices

## Qwen (Alibaba)

Alibaba ships the widest range of sizes of any provider here, from sub-cent flash models to flagship Max tiers.

| Model | API id | Input /1M | Cached /1M | Output /1M | Context | Source |
| --- | --- | --- | --- | --- | --- | --- |
Qwen3.7 Flash | `qwen3.7-flash` | $0.030 | $0.006 | $0.130 | 1M | via OpenRouter
Qwen3 30B A3B Instruct 2507 | `qwen3-30b-a3b-instruct-2507` | $0.048 | — | $0.193 | 262K | via OpenRouter
Qwen3.5-Flash | `qwen3.5-flash-02-23` | $0.065 | — | $0.260 | 1M | via OpenRouter
Qwen3 Coder 30B A3B Instruct | `qwen3-coder-30b-a3b-instruct` | $0.070 | — | $0.280 | 262K | via OpenRouter
Qwen3 32B | `qwen3-32b` | $0.080 | — | $0.280 | 131K | via OpenRouter
Qwen3 235B A22B Instruct 2507 | `qwen3-235b-a22b-2507` | $0.090 | — | $0.550 | 262K | via OpenRouter
Qwen3 Next 80B A3B Instruct | `qwen3-next-80b-a3b-instruct` | $0.090 | — | $1.10 | 262K | via OpenRouter
Qwen2.5 7B Instruct | `qwen-2.5-7b-instruct` | $0.100 | — | $0.200 | 33K | via OpenRouter
Qwen3.5-9B | `qwen3.5-9b` | $0.100 | — | $0.150 | 262K | via OpenRouter
Qwen3 VL 32B Instruct | `qwen3-vl-32b-instruct` | $0.104 | — | $0.416 | 131K | via OpenRouter
Qwen3 8B | `qwen3-8b` | $0.117 | — | $0.455 | 131K | via OpenRouter
Qwen3 VL 8B Instruct | `qwen3-vl-8b-instruct` | $0.117 | — | $0.455 | 262K | via OpenRouter
Qwen3 14B | `qwen3-14b` | $0.120 | — | $0.240 | 131K | via OpenRouter
Qwen3 30B A3B | `qwen3-30b-a3b` | $0.120 | — | $0.500 | 131K | via OpenRouter
Qwen3 Coder Next | `qwen3-coder-next` | $0.120 | $0.070 | $0.800 | 262K | via OpenRouter
Qwen3.5-35B-A3B | `qwen3.5-35b-a3b` | $0.140 | — | $1.00 | 262K | via OpenRouter
Qwen3 Next 80B A3B Thinking | `qwen3-next-80b-a3b-thinking` | $0.150 | — | $1.20 | 262K | via OpenRouter
Qwen3 VL 30B A3B Instruct | `qwen3-vl-30b-a3b-instruct` | $0.150 | — | $0.600 | 262K | via OpenRouter
Qwen3.6 35B A3B | `qwen3.6-35b-a3b` | $0.150 | $0.050 | $1.00 | 262K | via OpenRouter
Qwen3 VL 8B Thinking | `qwen3-vl-8b-thinking` | $0.180 | — | $2.10 | 131K | via OpenRouter
Qwen3.6 Flash | `qwen3.6-flash` | $0.188 | — | $1.13 | 1M | via OpenRouter
Qwen3 Coder Flash | `qwen3-coder-flash` | $0.195 | $0.039 | $0.975 | 1M | via OpenRouter
Qwen3.5-27B | `qwen3.5-27b` | $0.195 | — | $1.56 | 262K | via OpenRouter
Qwen3 30B A3B Thinking 2507 | `qwen3-30b-a3b-thinking-2507` | $0.200 | — | $2.40 | 82K | via OpenRouter
Qwen3 VL 30B A3B Thinking | `qwen3-vl-30b-a3b-thinking` | $0.200 | — | $2.40 | 262K | via OpenRouter
Qwen3 VL 235B A22B Instruct | `qwen3-vl-235b-a22b-instruct` | $0.210 | $0.100 | $1.90 | 262K | via OpenRouter
Qwen3 235B A22B Thinking 2507 | `qwen3-235b-a22b-thinking-2507` | $0.230 | — | $2.30 | 262K | via OpenRouter
Qwen2.5 VL 72B Instruct | `qwen2.5-vl-72b-instruct` | $0.250 | — | $0.750 | 128K | via OpenRouter
Qwen-Plus | `qwen-plus` | $0.260 | $0.052 | $0.780 | 1M | via OpenRouter
Qwen Plus 0728 | `qwen-plus-2025-07-28` | $0.260 | — | $0.780 | 1M | via OpenRouter
Qwen3.5 Plus 2026-02-15 | `qwen3.5-plus-02-15` | $0.260 | — | $1.56 | 1M | via OpenRouter
Qwen3.5-122B-A10B | `qwen3.5-122b-a10b` | $0.290 | — | $2.40 | 262K | via OpenRouter
Qwen3 Coder 480B A35B | `qwen3-coder` | $0.300 | $0.100 | $1.00 | 262K | via OpenRouter
Qwen3.5 Plus 2026-04-20 | `qwen3.5-plus-20260420` | $0.300 | — | $1.80 | 1M | via OpenRouter
Qwen3.7 Plus | `qwen3.7-plus` | $0.320 | $0.064 | $1.28 | 1M | via OpenRouter
Qwen3.6 Plus | `qwen3.6-plus` | $0.325 | — | $1.95 | 1M | via OpenRouter
Qwen2.5 72B Instruct | `qwen-2.5-72b-instruct` | $0.360 | — | $0.400 | 33K | via OpenRouter
Qwen Plus 0728 (thinking) | `qwen-plus-2025-07-28:thinking` | $0.400 | — | $1.20 | 1M | via OpenRouter
Qwen3 VL 235B A22B Thinking | `qwen3-vl-235b-a22b-thinking` | $0.400 | — | $4.00 | 131K | via OpenRouter
Qwen3 235B A22B | `qwen3-235b-a22b` | $0.455 | — | $1.82 | 131K | via OpenRouter
Qwen3.5 397B A17B | `qwen3.5-397b-a17b` | $0.500 | $0.300 | $3.60 | 262K | via OpenRouter
Qwen3.6 27B | `qwen3.6-27b` | $0.600 | $0.120 | $3.60 | 262K | via OpenRouter
Qwen3 Coder Plus | `qwen3-coder-plus` | $0.650 | $0.130 | $3.25 | 1M | via OpenRouter
Qwen2.5 Coder 32B Instruct | `qwen-2.5-coder-32b-instruct` | $0.660 | — | $1.00 | 33K | via OpenRouter
Qwen3 Max | `qwen3-max` | $0.780 | $0.156 | $3.90 | 262K | via OpenRouter
Qwen3 Max Thinking | `qwen3-max-thinking` | $0.780 | — | $3.90 | 262K | via OpenRouter
Qwen3.6 Max Preview | `qwen3.6-max-preview` | $1.03 | — | $6.16 | 262K | via OpenRouter
Qwen3.7 Max | `qwen3.7-max` | $1.48 | $0.295 | $4.42 | 1M | via OpenRouter
Qwen3.8 Max | `qwen3.8-max` | $2.00 | $0.250 | $6.00 | 1M | via OpenRouter

- Qwen3.7 Flash: https://www.costoftoken.com/models/alibaba/qwen3.7-flash
- Qwen3 30B A3B Instruct 2507: https://www.costoftoken.com/models/alibaba/qwen3-30b-a3b-instruct-2507
- Qwen3.5-Flash: https://www.costoftoken.com/models/alibaba/qwen3.5-flash-02-23
- Qwen3 Coder 30B A3B Instruct: https://www.costoftoken.com/models/alibaba/qwen3-coder-30b-a3b-instruct
- Qwen3 32B: https://www.costoftoken.com/models/alibaba/qwen3-32b
- Qwen3 235B A22B Instruct 2507: https://www.costoftoken.com/models/alibaba/qwen3-235b-a22b-2507
- Qwen3 Next 80B A3B Instruct: https://www.costoftoken.com/models/alibaba/qwen3-next-80b-a3b-instruct
- Qwen2.5 7B Instruct: https://www.costoftoken.com/models/alibaba/qwen-2.5-7b-instruct

## Claude (Anthropic)

Anthropic quotes prices per million tokens (MTok) and bills prompt caching as separate write and read rates, with cache hits far cheaper than base input.

| Model | API id | Input /1M | Cached /1M | Output /1M | Context | Source |
| --- | --- | --- | --- | --- | --- | --- |
Claude Haiku 3.5 | `claude-haiku-3.5` | $0.800 | $0.080 | $4.00 | — | first-party
Claude Haiku 4.5 | `claude-haiku-4.5` | $1.00 | $0.100 | $5.00 | 200K | first-party
Claude Sonnet 5 | `claude-sonnet-5` | $2.00 | $0.200 | $10.00 | 1M | first-party
Claude Sonnet 4 | `claude-sonnet-4` | $3.00 | $0.300 | $15.00 | 1M | first-party
Claude Sonnet 4.5 | `claude-sonnet-4.5` | $3.00 | $0.300 | $15.00 | 1M | first-party
Claude Sonnet 4.6 | `claude-sonnet-4.6` | $3.00 | $0.300 | $15.00 | 1M | first-party
Claude Opus 4.5 | `claude-opus-4.5` | $5.00 | $0.500 | $25.00 | 200K | first-party
Claude Opus 4.6 | `claude-opus-4.6` | $5.00 | $0.500 | $25.00 | 1M | first-party
Claude Opus 4.7 | `claude-opus-4.7` | $5.00 | $0.500 | $25.00 | 1M | first-party
Claude Opus 4.8 | `claude-opus-4.8` | $5.00 | $0.500 | $25.00 | 1M | first-party
Claude Opus 5 | `claude-opus-5` | $5.00 | $0.500 | $25.00 | 1M | first-party
Claude Fable 5 | `claude-fable-5` | $10.00 | $1.00 | $50.00 | 1M | first-party
Claude Mythos 5 | `claude-mythos-5` | $10.00 | $1.00 | $50.00 | — | first-party
Claude Opus 4 | `claude-opus-4` | $15.00 | $1.50 | $75.00 | 200K | first-party
Claude Opus 4.1 | `claude-opus-4.1` | $15.00 | $1.50 | $75.00 | 200K | first-party

- Claude Haiku 3.5: https://www.costoftoken.com/models/anthropic/claude-haiku-3.5
- Claude Haiku 4.5: https://www.costoftoken.com/models/anthropic/claude-haiku-4.5
- Claude Sonnet 5: https://www.costoftoken.com/models/anthropic/claude-sonnet-5
- Claude Sonnet 4: https://www.costoftoken.com/models/anthropic/claude-sonnet-4
- Claude Sonnet 4.5: https://www.costoftoken.com/models/anthropic/claude-sonnet-4.5
- Claude Sonnet 4.6: https://www.costoftoken.com/models/anthropic/claude-sonnet-4.6
- Claude Opus 4.5: https://www.costoftoken.com/models/anthropic/claude-opus-4.5
- Claude Opus 4.6: https://www.costoftoken.com/models/anthropic/claude-opus-4.6

## ERNIE (Baidu)

Baidu sells ERNIE through the Qianfan platform.

| Model | API id | Input /1M | Cached /1M | Output /1M | Context | Source |
| --- | --- | --- | --- | --- | --- | --- |
ERNIE 4.5 VL 424B A47B  | `ernie-4.5-vl-424b-a47b` | $0.420 | — | $1.25 | 123K | via OpenRouter

- ERNIE 4.5 VL 424B A47B : https://www.costoftoken.com/models/baidu/ernie-4.5-vl-424b-a47b

## Doubao (ByteDance)

ByteDance sells Doubao through Volcengine Ark, priced well below Western equivalents.

| Model | API id | Input /1M | Cached /1M | Output /1M | Context | Source |
| --- | --- | --- | --- | --- | --- | --- |
UI-TARS 7B  | `ui-tars-1.5-7b` | $0.100 | $0.100 | $0.200 | 128K | via OpenRouter

- UI-TARS 7B : https://www.costoftoken.com/models/bytedance/ui-tars-1.5-7b

## DeepSeek (DeepSeek)

DeepSeek is among the cheapest capable APIs, with aggressive cache-hit pricing that rewards repeated prompt prefixes.

| Model | API id | Input /1M | Cached /1M | Output /1M | Context | Source |
| --- | --- | --- | --- | --- | --- | --- |
DeepSeek V4 Flash Latest | `deepseek-v4-flash-latest` | $0.072 | $0.014 | $0.144 | 1M | via OpenRouter
DeepSeek V4 Flash 0731 | `deepseek-v4-flash-0731` | $0.080 | $0.016 | $0.180 | 1M | via OpenRouter
DeepSeek V4 Flash 0423 | `deepseek-v4-flash` | $0.140 | $0.028 | $0.280 | 1M | via OpenRouter
DeepSeek V3.1 | `deepseek-chat-v3.1` | $0.250 | $0.130 | $0.950 | 164K | via OpenRouter
DeepSeek V3 | `deepseek-chat` | $0.257 | — | $1.03 | 164K | via OpenRouter
DeepSeek V3.2 | `deepseek-v3.2` | $0.269 | $0.135 | $0.400 | 164K | via OpenRouter
DeepSeek V3 0324 | `deepseek-chat-v3-0324` | $0.270 | $0.135 | $1.12 | 164K | via OpenRouter
DeepSeek V3.1 Terminus | `deepseek-v3.1-terminus` | $0.270 | $0.135 | $1.00 | 164K | via OpenRouter
DeepSeek V3.2 Exp | `deepseek-v3.2-exp` | $0.270 | — | $0.410 | 164K | via OpenRouter
R1 0528 | `deepseek-r1-0528` | $0.500 | $0.350 | $2.15 | 164K | via OpenRouter
DeepSeek V4 Pro | `deepseek-v4-pro` | $0.632 | $0.053 | $1.26 | 1M | via OpenRouter
R1 | `deepseek-r1` | $0.700 | — | $2.50 | 164K | via OpenRouter
R1 Distill Llama 70B | `deepseek-r1-distill-llama-70b` | $0.800 | — | $0.800 | 8K | via OpenRouter

- DeepSeek V4 Flash Latest: https://www.costoftoken.com/models/deepseek/deepseek-v4-flash-latest
- DeepSeek V4 Flash 0731: https://www.costoftoken.com/models/deepseek/deepseek-v4-flash-0731
- DeepSeek V4 Flash 0423: https://www.costoftoken.com/models/deepseek/deepseek-v4-flash
- DeepSeek V3.1: https://www.costoftoken.com/models/deepseek/deepseek-chat-v3.1
- DeepSeek V3: https://www.costoftoken.com/models/deepseek/deepseek-chat
- DeepSeek V3.2: https://www.costoftoken.com/models/deepseek/deepseek-v3.2
- DeepSeek V3 0324: https://www.costoftoken.com/models/deepseek/deepseek-chat-v3-0324
- DeepSeek V3.1 Terminus: https://www.costoftoken.com/models/deepseek/deepseek-v3.1-terminus

## Gemini (Google)

Google publishes a free tier alongside paid pricing, and several Gemini models charge a higher rate above a long-context threshold.

| Model | API id | Input /1M | Cached /1M | Output /1M | Context | Source |
| --- | --- | --- | --- | --- | --- | --- |
Gemini 2.0 Flash-Lite | `gemini-2.0-flash-lite` | $0.075 | — | $0.300 | — | first-party
Gemini 2.0 Flash | `gemini-2.0-flash` | $0.100 | $0.025 | $0.400 | — | first-party
Gemini 2.5 Flash-Lite | `gemini-2.5-flash-lite` | $0.100 | $0.010 | $0.400 | 1M | first-party
Gemini 2.5 Flash-Lite Preview | `gemini-2.5-flash-lite-preview` | $0.100 | $0.010 | $0.400 | — | first-party
Gemini Embedding | `gemini-embedding` | $0.150 | — | — | — | first-party
Gemini Embedding 2 | `gemini-embedding-2` | $0.200 | — | — | — | first-party
Gemini 3.1 Flash-Lite | `gemini-3.1-flash-lite` | $0.250 | $0.025 | $1.50 | 1M | first-party
Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite) 🍌 | `gemini-3.1-flash-lite-image` | $0.250 | — | $1.50 | 66K | first-party
Gemini 2.5 Flash | `gemini-2.5-flash` | $0.300 | $0.030 | $2.50 | 1M | first-party
Gemini 2.5 Flash Image (Nano Banana) 🍌 | `gemini-2.5-flash-image` | $0.300 | — | $0.039 | 33K | first-party
Gemini 3.5 Flash-Lite | `gemini-3.5-flash-lite` | $0.300 | $0.030 | $2.50 | 1M | first-party
Gemini 2.5 Flash Native Audio (Live API) | `gemini-2.5-flash-native-audio` | $0.500 | — | $2.00 | — | first-party
Gemini 2.5 Flash Preview TTS | `gemini-2.5-flash-preview-tts` | $0.500 | — | $10.00 | — | first-party
Gemini 3 Flash Preview | `gemini-3-flash-preview` | $0.500 | $0.050 | $3.00 | 1M | first-party
Gemini 3.1 Flash Image (Nano Banana 2) 🍌 | `gemini-3.1-flash-image` | $0.500 | — | $3.00 | 131K | first-party
Gemini 3.1 Flash Live Preview | `gemini-3.1-flash-live-preview` | $0.750 | — | $4.50 | — | first-party
Gemini 2.5 Pro Preview TTS | `gemini-2.5-pro-preview-tts` | $1.00 | — | $20.00 | — | first-party
Gemini 3.1 Flash TTS Preview | `gemini-3.1-flash-tts-preview` | $1.00 | — | $20.00 | — | first-party
Gemini Robotics ER 1.6 Preview | `gemini-robotics-er-1.6-preview` | $1.00 | — | $5.00 | — | first-party
Gemini 2.5 Computer Use Preview | `gemini-2.5-computer-use-preview` | $1.25 | — | $10.00 | — | first-party
Gemini 2.5 Pro | `gemini-2.5-pro` | $1.25 | $0.125 | $10.00 | 1M | first-party
Gemini 3.5 Flash | `gemini-3.5-flash` | $1.50 | $0.150 | $9.00 | 1M | first-party
Gemini 3.6 Flash | `gemini-3.6-flash` | $1.50 | $0.150 | $7.50 | 1M | first-party
Gemini Omni Flash Preview | `gemini-omni-flash-preview` | $1.50 | — | $9.00 | — | first-party
Gemini 3 Pro Image (Nano Banana Pro) 🍌 | `gemini-3-pro-image` | $2.00 | — | $12.00 | 131K | first-party
Gemini 3.1 Pro Preview | `gemini-3.1-pro-preview` | $2.00 | $0.200 | $12.00 | 1M | first-party
Gemini Robotics ER 2 Preview | `gemini-robotics-er-2-preview` | $2.00 | $0.200 | $10.00 | — | first-party
Gemini Robotics ER 2 Streaming Preview | `gemini-robotics-er-2-streaming-preview` | $2.00 | — | $10.00 | — | first-party
Gemini 3.5 Live Translate | `gemini-3.5-live-translate` | $3.50 | — | $21.00 | — | first-party

Long-context tiers:

- Gemini 2.5 Computer Use Preview: above 200K tokens, input $2.50 and output $15.00 per 1M.
- Gemini 2.5 Pro: above 200K tokens, input $2.50 and output $15.00 per 1M.
- Gemini 3.1 Pro Preview: above 200K tokens, input $4.00 and output $18.00 per 1M.

- Gemini 2.0 Flash-Lite: https://www.costoftoken.com/models/google/gemini-2.0-flash-lite
- Gemini 2.0 Flash: https://www.costoftoken.com/models/google/gemini-2.0-flash
- Gemini 2.5 Flash-Lite: https://www.costoftoken.com/models/google/gemini-2.5-flash-lite
- Gemini 2.5 Flash-Lite Preview: https://www.costoftoken.com/models/google/gemini-2.5-flash-lite-preview
- Gemini Embedding: https://www.costoftoken.com/models/google/gemini-embedding
- Gemini Embedding 2: https://www.costoftoken.com/models/google/gemini-embedding-2
- Gemini 3.1 Flash-Lite: https://www.costoftoken.com/models/google/gemini-3.1-flash-lite
- Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite) 🍌: https://www.costoftoken.com/models/google/gemini-3.1-flash-lite-image

## Kimi (Moonshot AI)

Moonshot’s Kimi models target long-context work at a fraction of Western flagship pricing.

| Model | API id | Input /1M | Cached /1M | Output /1M | Context | Source |
| --- | --- | --- | --- | --- | --- | --- |
Kimi K2.7 Code (batch) | `kimi-k2.7-code:batch` | $0.475 | $0.095 | $2.00 | 262K | via OpenRouter
Kimi K2 0711 | `kimi-k2` | $0.570 | — | $2.30 | 131K | via OpenRouter
Kimi K2.5 | `kimi-k2.5` | $0.570 | $0.095 | $2.85 | 262K | via OpenRouter
Kimi K2.6 | `kimi-k2.6` | $0.580 | $0.098 | $2.44 | 262K | via OpenRouter
Kimi K2 0905 | `kimi-k2-0905` | $0.600 | — | $2.50 | 262K | via OpenRouter
Kimi K2 Thinking | `kimi-k2-thinking` | $0.600 | $0.150 | $2.50 | 262K | via OpenRouter
Kimi K2.7 Code | `kimi-k2.7-code` | $0.700 | $0.150 | $3.50 | 262K | via OpenRouter
MoonshotAI Kimi Latest | `kimi-latest` | $2.80 | $0.290 | $14.00 | 1M | via OpenRouter
Kimi K3 | `kimi-k3` | $3.00 | $0.300 | $15.00 | 1M | via OpenRouter

- Kimi K2.7 Code (batch): https://www.costoftoken.com/models/moonshot/kimi-k2.7-code%3Abatch
- Kimi K2 0711: https://www.costoftoken.com/models/moonshot/kimi-k2
- Kimi K2.5: https://www.costoftoken.com/models/moonshot/kimi-k2.5
- Kimi K2.6: https://www.costoftoken.com/models/moonshot/kimi-k2.6
- Kimi K2 0905: https://www.costoftoken.com/models/moonshot/kimi-k2-0905
- Kimi K2 Thinking: https://www.costoftoken.com/models/moonshot/kimi-k2-thinking
- Kimi K2.7 Code: https://www.costoftoken.com/models/moonshot/kimi-k2.7-code
- MoonshotAI Kimi Latest: https://www.costoftoken.com/models/moonshot/kimi-latest

## OpenAI (OpenAI)

OpenAI prices most models in short- and long-context tiers, and offers cached input at a large discount for repeated prompt prefixes.

| Model | API id | Input /1M | Cached /1M | Output /1M | Context | Source |
| --- | --- | --- | --- | --- | --- | --- |
omni-moderation-latest | `omni-moderation-latest` | Free | — | — | — | first-party
text-embedding-3-small | `text-embedding-3-small` | $0.020 | — | — | — | first-party
gpt-5-nano | `gpt-5-nano` | $0.050 | $0.005 | $0.400 | 400K | first-party
gpt-4.1-nano | `gpt-4.1-nano` | $0.100 | $0.025 | $0.400 | 1M | first-party
text-embedding-ada-002 | `text-embedding-ada-002` | $0.100 | — | — | — | first-party
text-embedding-3-large | `text-embedding-3-large` | $0.130 | — | — | — | first-party
gpt-4o-mini | `gpt-4o-mini` | $0.150 | $0.075 | $0.600 | 128K | first-party
gpt-4.1-nano-2025-04-14 | `gpt-4.1-nano-2025-04-14` | $0.200 | $0.050 | $0.800 | — | first-party
gpt-5.4-nano | `gpt-5.4-nano` | $0.200 | $0.020 | $1.25 | 400K | first-party
gpt-5.6-luna | `gpt-5.6-luna` | $0.200 | $0.020 | $1.20 | 1.1M | first-party
gpt-5-mini | `gpt-5-mini` | $0.250 | $0.025 | $2.00 | 400K | first-party
gpt-4o-mini-2024-07-18 | `gpt-4o-mini-2024-07-18` | $0.300 | $0.150 | $1.20 | 128K | first-party
babbage-002 | `babbage-002` | $0.400 | — | $0.400 | — | first-party
gpt-4.1-mini | `gpt-4.1-mini` | $0.400 | $0.100 | $1.60 | 1M | first-party
gpt-3.5-turbo | `gpt-3.5-turbo` | $0.500 | — | $1.50 | 16K | first-party
gpt-3.5-turbo-0125 | `gpt-3.5-turbo-0125` | $0.500 | — | $1.50 | — | first-party
gpt-4o-mini-tts | `gpt-4o-mini-tts` | $0.600 | — | — | — | first-party
gpt-5.4-mini | `gpt-5.4-mini` | $0.750 | $0.075 | $4.50 | 400K | first-party
gpt-4.1-mini-2025-04-14 | `gpt-4.1-mini-2025-04-14` | $0.800 | $0.200 | $3.20 | — | first-party
gpt-3.5-turbo-1106 | `gpt-3.5-turbo-1106` | $1.00 | — | $2.00 | — | first-party
o3-mini | `o3-mini` | $1.10 | $0.550 | $4.40 | 200K | first-party
o4-mini | `o4-mini` | $1.10 | $0.275 | $4.40 | 200K | first-party
gpt-4o-mini-transcribe | `gpt-4o-mini-transcribe` | $1.25 | — | $5.00 | — | first-party
gpt-5 | `gpt-5` | $1.25 | $0.125 | $10.00 | 400K | first-party
gpt-5-search-api | `gpt-5-search-api` | $1.25 | $0.125 | $10.00 | — | first-party
gpt-5.1 | `gpt-5.1` | $1.25 | $0.125 | $10.00 | 400K | first-party
gpt-3.5-turbo-instruct | `gpt-3.5-turbo-instruct` | $1.50 | — | $2.00 | 4K | first-party
gpt-5.2 | `gpt-5.2` | $1.75 | $0.175 | $14.00 | 400K | first-party
gpt-5.2-chat-latest | `gpt-5.2-chat-latest` | $1.75 | $0.175 | $14.00 | — | first-party
gpt-5.3-chat-latest | `gpt-5.3-chat-latest` | $1.75 | $0.175 | $14.00 | — | first-party
gpt-5.3-codex | `gpt-5.3-codex` | $1.75 | $0.175 | $14.00 | 400K | first-party
davinci-002 | `davinci-002` | $2.00 | — | $2.00 | — | first-party
gpt-4.1 | `gpt-4.1` | $2.00 | $0.500 | $8.00 | 1M | first-party
gpt-5.6-terra | `gpt-5.6-terra` | $2.00 | $0.200 | $12.00 | 1.1M | first-party
o3 | `o3` | $2.00 | $0.500 | $8.00 | 200K | first-party
gpt-4o | `gpt-4o` | $2.50 | $1.25 | $10.00 | 128K | first-party
gpt-4o-transcribe | `gpt-4o-transcribe` | $2.50 | — | $10.00 | — | first-party
gpt-4o-transcribe-diarize | `gpt-4o-transcribe-diarize` | $2.50 | — | $10.00 | — | first-party
gpt-5.4 | `gpt-5.4` | $2.50 | $0.250 | $15.00 | 1.1M | first-party
gpt-image-1-mini | `gpt-image-1-mini` | $2.50 | $0.250 | $8.00 | — | first-party
gpt-4.1-2025-04-14 | `gpt-4.1-2025-04-14` | $3.00 | $0.750 | $12.00 | — | first-party
gpt-4o-2024-08-06 | `gpt-4o-2024-08-06` | $3.75 | $1.88 | $15.00 | 128K | first-party
o4-mini-2025-04-16 | `o4-mini-2025-04-16` | $4.00 | $1.00 | $16.00 | — | first-party
chat-latest | `chat-latest` | $5.00 | $0.500 | $30.00 | — | first-party
gpt-4o-2024-05-13 | `gpt-4o-2024-05-13` | $5.00 | — | $15.00 | 128K | first-party
gpt-5.5 | `gpt-5.5` | $5.00 | $0.500 | $30.00 | 1.1M | first-party
gpt-5.6-sol | `gpt-5.6-sol` | $5.00 | $0.500 | $30.00 | 1.1M | first-party
chatgpt-image-latest | `chatgpt-image-latest` | $8.00 | $2.00 | $32.00 | — | first-party
gpt-image-1.5 | `gpt-image-1.5` | $8.00 | $2.00 | $32.00 | — | first-party
gpt-image-2 | `gpt-image-2` | $8.00 | $2.00 | $30.00 | — | first-party
gpt-4-turbo-2024-04-09 | `gpt-4-turbo-2024-04-09` | $10.00 | — | $30.00 | — | first-party
gpt-audio-mini | `gpt-audio-mini` | $10.00 | — | — | 128K | first-party
gpt-image-1 | `gpt-image-1` | $10.00 | $2.50 | $40.00 | — | first-party
gpt-realtime-2.1-mini | `gpt-realtime-2.1-mini` | $10.00 | $0.300 | — | — | first-party
gpt-realtime-mini | `gpt-realtime-mini` | $10.00 | $0.300 | — | — | first-party
gpt-5.5-cyber | `gpt-5.5-cyber` | $12.50 | $1.25 | $75.00 | — | first-party
gpt-5.6-cyber | `gpt-5.6-cyber` | $12.50 | $1.25 | $75.00 | — | first-party
gpt-5-pro | `gpt-5-pro` | $15.00 | — | $120.00 | 400K | first-party
o1 | `o1` | $15.00 | $7.50 | $60.00 | 200K | first-party
tts-1 | `tts-1` | $15.00 | — | — | — | first-party
o3-pro | `o3-pro` | $20.00 | — | $80.00 | 200K | first-party
gpt-5.2-pro | `gpt-5.2-pro` | $21.00 | — | $168.00 | 400K | first-party
gpt-4-0613 | `gpt-4-0613` | $30.00 | — | $60.00 | — | first-party
gpt-5.4-pro | `gpt-5.4-pro` | $30.00 | — | $180.00 | 1.1M | first-party
gpt-5.5-pro | `gpt-5.5-pro` | $30.00 | — | $180.00 | 1.1M | first-party
tts-1-hd | `tts-1-hd` | $30.00 | — | — | — | first-party
gpt-audio | `gpt-audio` | $32.00 | — | — | 128K | first-party
gpt-audio-1.5 | `gpt-audio-1.5` | $32.00 | — | — | — | first-party
gpt-realtime | `gpt-realtime` | $32.00 | $0.400 | — | — | first-party
gpt-realtime-1.5 | `gpt-realtime-1.5` | $32.00 | $0.400 | — | — | first-party
gpt-realtime-2 | `gpt-realtime-2` | $32.00 | $0.400 | — | — | first-party
gpt-realtime-2.1 | `gpt-realtime-2.1` | $32.00 | $0.400 | — | — | first-party
o1-pro | `o1-pro` | $150.00 | — | $600.00 | 200K | first-party

Long-context tiers:

- gpt-5.6-luna: above 128K tokens, input $0.400 and output $1.80 per 1M.
- gpt-5.6-terra: above 128K tokens, input $4.00 and output $18.00 per 1M.
- gpt-5.4: above 272K tokens, input $5.00 and output $22.50 per 1M.
- gpt-5.5: above 272K tokens, input $10.00 and output $45.00 per 1M.
- gpt-5.6-sol: above 128K tokens, input $10.00 and output $45.00 per 1M.
- gpt-5.4-pro: above 272K tokens, input $60.00 and output $270.00 per 1M.
- gpt-5.5-pro: above 272K tokens, input $60.00 and output $270.00 per 1M.

- omni-moderation-latest: https://www.costoftoken.com/models/openai/omni-moderation-latest
- text-embedding-3-small: https://www.costoftoken.com/models/openai/text-embedding-3-small
- gpt-5-nano: https://www.costoftoken.com/models/openai/gpt-5-nano
- gpt-4.1-nano: https://www.costoftoken.com/models/openai/gpt-4.1-nano
- text-embedding-ada-002: https://www.costoftoken.com/models/openai/text-embedding-ada-002
- text-embedding-3-large: https://www.costoftoken.com/models/openai/text-embedding-3-large
- gpt-4o-mini: https://www.costoftoken.com/models/openai/gpt-4o-mini
- gpt-4.1-nano-2025-04-14: https://www.costoftoken.com/models/openai/gpt-4.1-nano-2025-04-14

## Grok (xAI)

xAI publishes a machine-readable model catalogue, including context window and a long-context tier that applies above 128K tokens.

| Model | API id | Input /1M | Cached /1M | Output /1M | Context | Source |
| --- | --- | --- | --- | --- | --- | --- |
grok-build-0.1 | `grok-build-0.1` | $1.00 | $0.200 | $2.00 | 256K | first-party
grok-4.20-0309-non-reasoning | `grok-4.20-0309-non-reasoning` | $1.25 | $0.200 | $2.50 | 1M | first-party
grok-4.20-0309-reasoning | `grok-4.20-0309-reasoning` | $1.25 | $0.200 | $2.50 | 1M | first-party
grok-4.20-multi-agent-0309 | `grok-4.20-multi-agent-0309` | $1.25 | $0.200 | $2.50 | 1M | first-party
grok-4.3 | `grok-4.3` | $1.25 | $0.200 | $2.50 | 1M | first-party
grok-4.5 | `grok-4.5` | $2.00 | $0.300 | $6.00 | 500K | first-party

Long-context tiers:

- grok-build-0.1: above 128K tokens, input $2.00 and output $4.00 per 1M.
- grok-4.20-0309-non-reasoning: above 128K tokens, input $2.50 and output $5.00 per 1M.
- grok-4.20-0309-reasoning: above 128K tokens, input $2.50 and output $5.00 per 1M.
- grok-4.20-multi-agent-0309: above 128K tokens, input $2.50 and output $5.00 per 1M.
- grok-4.3: above 128K tokens, input $2.50 and output $5.00 per 1M.
- grok-4.5: above 128K tokens, input $4.00 and output $12.00 per 1M.

- grok-build-0.1: https://www.costoftoken.com/models/xai/grok-build-0.1
- grok-4.20-0309-non-reasoning: https://www.costoftoken.com/models/xai/grok-4.20-0309-non-reasoning
- grok-4.20-0309-reasoning: https://www.costoftoken.com/models/xai/grok-4.20-0309-reasoning
- grok-4.20-multi-agent-0309: https://www.costoftoken.com/models/xai/grok-4.20-multi-agent-0309
- grok-4.3: https://www.costoftoken.com/models/xai/grok-4.3
- grok-4.5: https://www.costoftoken.com/models/xai/grok-4.5

## GLM (Zhipu AI)

Zhipu publishes several genuinely free Flash models alongside paid GLM tiers, which is unusual among hosted APIs.

| Model | API id | Input /1M | Cached /1M | Output /1M | Context | Source |
| --- | --- | --- | --- | --- | --- | --- |
GLM-4.5-Flash | `glm-4.5-flash` | Free | Free | Free | — | first-party
GLM-4.6V-Flash | `glm-4.6v-flash` | Free | Free | Free | — | first-party
GLM-4.7-Flash | `glm-4.7-flash` | Free | Free | Free | 203K | first-party
GLM-OCR | `glm-ocr` | $0.030 | — | $0.030 | — | first-party
GLM-4.6V-FlashX | `glm-4.6v-flashx` | $0.040 | $0.004 | $0.400 | — | first-party
GLM-4.7-FlashX | `glm-4.7-flashx` | $0.070 | $0.010 | $0.400 | — | first-party
GLM-4-32B-0414-128K | `glm-4-32b-0414-128k` | $0.100 | — | $0.100 | — | first-party
GLM-4.5-Air | `glm-4.5-air` | $0.200 | $0.030 | $1.10 | 131K | first-party
GLM-4.6V | `glm-4.6v` | $0.300 | $0.050 | $0.900 | 131K | first-party
GLM-4.5 | `glm-4.5` | $0.600 | $0.110 | $2.20 | 131K | first-party
GLM-4.5V | `glm-4.5v` | $0.600 | $0.110 | $1.80 | 66K | first-party
GLM-4.6 | `glm-4.6` | $0.600 | $0.110 | $2.20 | 205K | first-party
GLM-4.7 | `glm-4.7` | $0.600 | $0.110 | $2.20 | 205K | first-party
GLM-5 | `glm-5` | $1.00 | $0.200 | $3.20 | 205K | first-party
GLM-4.5-AirX | `glm-4.5-airx` | $1.10 | $0.220 | $4.50 | — | first-party
GLM-5-Turbo | `glm-5-turbo` | $1.20 | $0.240 | $4.00 | 203K | first-party
GLM-5V-Turbo | `glm-5v-turbo` | $1.20 | $0.240 | $4.00 | 203K | first-party
GLM-5.1 | `glm-5.1` | $1.40 | $0.260 | $4.40 | 205K | first-party
GLM-5.2 | `glm-5.2` | $1.40 | $0.260 | $4.40 | 1M | first-party
GLM-4.5-X | `glm-4.5-x` | $2.20 | $0.450 | $8.90 | — | first-party

- GLM-4.5-Flash: https://www.costoftoken.com/models/zhipu/glm-4.5-flash
- GLM-4.6V-Flash: https://www.costoftoken.com/models/zhipu/glm-4.6v-flash
- GLM-4.7-Flash: https://www.costoftoken.com/models/zhipu/glm-4.7-flash
- GLM-OCR: https://www.costoftoken.com/models/zhipu/glm-ocr
- GLM-4.6V-FlashX: https://www.costoftoken.com/models/zhipu/glm-4.6v-flashx
- GLM-4.7-FlashX: https://www.costoftoken.com/models/zhipu/glm-4.7-flashx
- GLM-4-32B-0414-128K: https://www.costoftoken.com/models/zhipu/glm-4-32b-0414-128k
- GLM-4.5-Air: https://www.costoftoken.com/models/zhipu/glm-4.5-air

## Citing this data

Include the last-updated date. Prices change frequently and a quoted figure
without a date is unverifiable. Canonical source: https://www.costoftoken.com/
