LLM cost calculator

Describe the requests you actually send and every model is priced against them, cheapest first. Output typically costs several times more than input, so the cheapest model for a summariser is rarely the cheapest for a chat agent — which a per-token price list cannot tell you.

Start from a typical workload

Prompt, system message and any retrieved context

What the model generates back

Total calls across all users

Share of the prompt that repeats between calls. Usually the single largest saving available.

$6.30
Cheapest — GLM-OCR
$58.5K
Most expensive — o1-pro
9286×
Difference between them

3 models would cost nothing

Listed separately rather than at the top of the ranking, where they would win every comparison by default and tell you nothing: GLM-4.5-Flash, GLM-4.6V-Flash, GLM-4.7-Flash.

Monthly cost, cheapest first

195 paid models priced for 100,000 requests of 1,500 in / 600 out.

Models ranked by estimated monthly cost
#ModelProviderPer requestPer monthContext
1GLM-OCRZhipu AI (GLM)$0.0001$6.30
2Qwen3.7 FlashVia OpenRouterAlibaba (Qwen)$0.0001$11.221M
3DeepSeek V4 Flash LatestVia OpenRouterDeepSeek$0.0002$16.851M
4Qwen3 30B A3B Instruct 2507Via OpenRouterAlibaba (Qwen)$0.0002$18.81262K
5DeepSeek V4 Flash 0731Via OpenRouterDeepSeek$0.0002$19.921M
6GLM-4-32B-0414-128KZhipu AI (GLM)$0.0002$21.00
7Qwen3.5-9BVia OpenRouterAlibaba (Qwen)$0.0002$24.00262K
8Qwen3.5-FlashVia OpenRouterAlibaba (Qwen)$0.0003$25.351M
9Qwen2.5 7B InstructVia OpenRouterAlibaba (Qwen)$0.0003$27.0033K
10UI-TARS 7B Via OpenRouterByteDance (Doubao)$0.0003$27.00128K
11Qwen3 Coder 30B A3B InstructVia OpenRouterAlibaba (Qwen)$0.0003$27.30262K
12GLM-4.6V-FlashXZhipu AI (GLM)$0.0003$28.38
13Qwen3 32BVia OpenRouterAlibaba (Qwen)$0.0003$28.80131K
14Gemini 2.0 Flash-LiteGoogle$0.0003$29.25
15gpt-5-nanoOpenAI$0.0003$29.47400K
16GLM-4.7-FlashXZhipu AI (GLM)$0.0003$31.80
17Qwen3 14BVia OpenRouterAlibaba (Qwen)$0.0003$32.40131K
18DeepSeek V4 Flash 0423Via OpenRouterDeepSeek$0.0003$32.761M
19Gemini 2.5 Flash-LiteGoogle$0.0003$34.951M
20Gemini 2.5 Flash-Lite PreviewGoogle$0.0003$34.95
21Gemini 2.0 FlashGoogle$0.0004$35.63
22gpt-4.1-nanoOpenAI$0.0004$35.631M
23Qwen3 VL 32B InstructVia OpenRouterAlibaba (Qwen)$0.0004$40.56131K
24Qwen3 8BVia OpenRouterAlibaba (Qwen)$0.0004$44.85131K
25Qwen3 VL 8B InstructVia OpenRouterAlibaba (Qwen)$0.0004$44.85262K
26Qwen3 235B A22B Instruct 2507Via OpenRouterAlibaba (Qwen)$0.0005$46.50262K
27Gemini 2.5 Flash Image (Nano Banana) 🍌Google$0.0005$47.3433K
28Qwen3 30B A3BVia OpenRouterAlibaba (Qwen)$0.0005$48.00131K
29gpt-4o-miniOpenAI$0.0006$55.13128K
30DeepSeek V3.2Via OpenRouterDeepSeek$0.0006$58.30164K
31Qwen3 VL 30B A3B InstructVia OpenRouterAlibaba (Qwen)$0.0006$58.50262K
32Qwen3 Coder NextVia OpenRouterAlibaba (Qwen)$0.0006$63.75262K
33DeepSeek V3.2 ExpVia OpenRouterDeepSeek$0.0007$65.10164K
34gpt-4.1-nano-2025-04-14OpenAI$0.0007$71.25
35Qwen-PlusVia OpenRouterAlibaba (Qwen)$0.0008$76.441M
36Qwen3.6 35B A3BVia OpenRouterAlibaba (Qwen)$0.0008$78.00262K
37Qwen2.5 72B InstructVia OpenRouterAlibaba (Qwen)$0.0008$78.0033K
38Qwen3 Next 80B A3B InstructVia OpenRouterAlibaba (Qwen)$0.0008$79.50262K
39Qwen3 Coder FlashVia OpenRouterAlibaba (Qwen)$0.0008$80.731M
40Qwen3.5-35B-A3BVia OpenRouterAlibaba (Qwen)$0.0008$81.00262K
41Qwen2.5 VL 72B InstructVia OpenRouterAlibaba (Qwen)$0.0008$82.50128K
42babbage-002OpenAI$0.0008$84.00
43Qwen Plus 0728Via OpenRouterAlibaba (Qwen)$0.0009$85.801M
44GLM-4.6VZhipu AI (GLM)$0.0009$87.75131K
45GLM-4.5-AirZhipu AI (GLM)$0.0009$88.35131K
46DeepSeek V3.1Via OpenRouterDeepSeek$0.0009$89.10164K
47gpt-5.6-lunaOpenAI$0.0009$93.901.1M
48DeepSeek V3.1 TerminusVia OpenRouterDeepSeek$0.0009$94.42164K
49Qwen3 Next 80B A3B ThinkingVia OpenRouterAlibaba (Qwen)$0.0009$94.50262K
50Qwen3.6 FlashVia OpenRouterAlibaba (Qwen)$0.0010$95.631M
51Qwen3 Coder 480B A35BVia OpenRouterAlibaba (Qwen)$0.0010$96.00262K
52gpt-5.4-nanoOpenAI$0.0010$96.90400K
53DeepSeek V3Via OpenRouterDeepSeek$0.0010$100.33164K
54DeepSeek V3 0324Via OpenRouterDeepSeek$0.0010$101.63164K
55gpt-4o-mini-2024-07-18OpenAI$0.0011$110.25128K
56Qwen3.7 PlusVia OpenRouterAlibaba (Qwen)$0.0011$113.281M
57Gemini 3.1 Flash-LiteGoogle$0.0012$117.381M
58Qwen3.5-27BVia OpenRouterAlibaba (Qwen)$0.0012$122.85262K
59Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite) 🍌Google$0.0013$127.5066K
60Qwen Plus 0728 (thinking)Via OpenRouterAlibaba (Qwen)$0.0013$132.001M
61Qwen3.5 Plus 2026-02-15Via OpenRouterAlibaba (Qwen)$0.0013$132.601M
62ERNIE 4.5 VL 424B A47B Via OpenRouterBaidu (ERNIE)$0.0014$138.00123K
63Qwen3 VL 235B A22B InstructVia OpenRouterAlibaba (Qwen)$0.0014$140.55262K
64gpt-4.1-miniOpenAI$0.0014$142.501M
65DeepSeek V4 ProVia OpenRouterDeepSeek$0.0014$144.531M
66gpt-5-miniOpenAI$0.0015$147.38400K
67Qwen3 VL 8B ThinkingVia OpenRouterAlibaba (Qwen)$0.0015$153.00131K
68Qwen3.5 Plus 2026-04-20Via OpenRouterAlibaba (Qwen)$0.0015$153.001M
69Qwen2.5 Coder 32B InstructVia OpenRouterAlibaba (Qwen)$0.0016$159.0033K
70gpt-3.5-turboOpenAI$0.0016$165.0016K
71gpt-3.5-turbo-0125OpenAI$0.0016$165.00
72Qwen3.6 PlusVia OpenRouterAlibaba (Qwen)$0.0017$165.751M
73R1 Distill Llama 70BVia OpenRouterDeepSeek$0.0017$168.008K
74Qwen3 235B A22B Thinking 2507Via OpenRouterAlibaba (Qwen)$0.0017$172.50262K
75Qwen3 30B A3B Thinking 2507Via OpenRouterAlibaba (Qwen)$0.0017$174.0082K
76Qwen3 VL 30B A3B ThinkingVia OpenRouterAlibaba (Qwen)$0.0017$174.00262K
77Kimi K2.7 Code (batch)Via OpenRouterMoonshot AI (Kimi)$0.0017$174.15262K
78GLM-4.5VZhipu AI (GLM)$0.0018$175.9566K
79Qwen3 235B A22BVia OpenRouterAlibaba (Qwen)$0.0018$177.45131K
80Gemini 2.5 FlashGoogle$0.0018$182.851M
81Gemini 3.5 Flash-LiteGoogle$0.0018$182.851M
82Qwen3.5-122B-A10BVia OpenRouterAlibaba (Qwen)$0.0019$187.50262K
83Gemini 2.5 Flash Native Audio (Live API)Google$0.0019$195.00
84R1 0528Via OpenRouterDeepSeek$0.0020$197.25164K
85GLM-4.5Zhipu AI (GLM)$0.0020$199.95131K
86GLM-4.6Zhipu AI (GLM)$0.0020$199.95205K
87GLM-4.7Zhipu AI (GLM)$0.0020$199.95205K
88Kimi K2.6Via OpenRouterMoonshot AI (Kimi)$0.0021$211.64262K
89Kimi K2 ThinkingVia OpenRouterMoonshot AI (Kimi)$0.0022$219.75262K
90Kimi K2 0711Via OpenRouterMoonshot AI (Kimi)$0.0022$223.50131K
91grok-build-0.1xAI$0.0023$234.00256K
92Gemini 3 Flash PreviewGoogle$0.0023$234.751M
93Kimi K2.5Via OpenRouterMoonshot AI (Kimi)$0.0024$235.13262K
94Kimi K2 0905Via OpenRouterMoonshot AI (Kimi)$0.0024$240.00262K
95Gemini 3.1 Flash Image (Nano Banana 2) 🍌Google$0.0026$255.00131K
96R1Via OpenRouterDeepSeek$0.0026$255.00164K
97Qwen3 Coder PlusVia OpenRouterAlibaba (Qwen)$0.0027$269.101M
98gpt-3.5-turbo-1106OpenAI$0.0027$270.00
99Qwen3.5 397B A17BVia OpenRouterAlibaba (Qwen)$0.0028$282.00262K
100Qwen3.6 27BVia OpenRouterAlibaba (Qwen)$0.0028$284.40262K
101gpt-4.1-mini-2025-04-14OpenAI$0.0029$285.00
102Kimi K2.7 CodeVia OpenRouterMoonshot AI (Kimi)$0.0029$290.25262K
103grok-4.20-0309-non-reasoningxAI$0.0029$290.251M
104grok-4.20-0309-reasoningxAI$0.0029$290.251M
105grok-4.20-multi-agent-0309xAI$0.0029$290.251M
106grok-4.3xAI$0.0029$290.251M
107Qwen3 VL 235B A22B ThinkingVia OpenRouterAlibaba (Qwen)$0.0030$300.00131K
108GLM-5Zhipu AI (GLM)$0.0031$306.00205K
109Qwen3 MaxVia OpenRouterAlibaba (Qwen)$0.0032$322.92262K
110Claude Haiku 3.5Anthropic$0.0033$327.60
111gpt-3.5-turbo-instructOpenAI$0.0034$345.004K
112Qwen3 Max ThinkingVia OpenRouterAlibaba (Qwen)$0.0035$351.00262K
113gpt-5.4-miniOpenAI$0.0035$352.13400K
114GLM-5-TurboZhipu AI (GLM)$0.0038$376.80203K
115GLM-5V-TurboZhipu AI (GLM)$0.0038$376.80203K
116Gemini 3.1 Flash Live PreviewGoogle$0.0038$382.50
117o4-miniOpenAI$0.0039$391.88200K
118GLM-4.5-AirXZhipu AI (GLM)$0.0040$395.40
119o3-miniOpenAI$0.0040$404.25200K
120Claude Haiku 4.5Anthropic$0.0041$409.50200K
121davinci-002OpenAI$0.0042$420.00
122GLM-5.1Zhipu AI (GLM)$0.0042$422.70205K
123GLM-5.2Zhipu AI (GLM)$0.0042$422.701M
124Qwen3.7 MaxVia OpenRouterAlibaba (Qwen)$0.0043$433.651M
125Gemini Robotics ER 1.6 PreviewGoogle$0.0045$450.00
126gpt-4o-mini-transcribeOpenAI$0.0049$487.50
127Qwen3.6 Max PreviewVia OpenRouterAlibaba (Qwen)$0.0052$523.77262K
128Qwen3.8 MaxVia OpenRouterAlibaba (Qwen)$0.0058$581.251M
129grok-4.5xAI$0.0058$583.50500K
130Gemini 3.6 FlashGoogle$0.0061$614.251M
131Gemini 2.5 Flash Preview TTSGoogle$0.0067$675.00
132Gemini 3.5 FlashGoogle$0.0070$704.251M
133gpt-4.1OpenAI$0.0071$712.501M
134o3OpenAI$0.0071$712.50200K
135Gemini 2.5 ProGoogle$0.0074$736.881M
136gpt-5OpenAI$0.0074$736.88400K
137gpt-5-search-apiOpenAI$0.0074$736.88
138gpt-5.1OpenAI$0.0074$736.88400K
139gpt-image-1-miniOpenAI$0.0075$753.75
140Gemini Omni Flash PreviewGoogle$0.0076$765.00
141GLM-4.5-XZhipu AI (GLM)$0.0079$785.25
142Gemini 2.5 Computer Use PreviewGoogle$0.0079$787.50
143Claude Sonnet 5Anthropic$0.0082$819.001M
144Gemini Robotics ER 2 PreviewGoogle$0.0082$819.00
145Gemini Robotics ER 2 Streaming PreviewGoogle$0.0090$900.00
146gpt-4oOpenAI$0.0092$918.75128K
147Gemini 3.1 Pro PreviewGoogle$0.0094$939.001M
148gpt-5.6-terraOpenAI$0.0094$939.001.1M
149gpt-4o-transcribeOpenAI$0.0097$975.00
150gpt-4o-transcribe-diarizeOpenAI$0.0097$975.00
151Gemini 3 Pro Image (Nano Banana Pro) 🍌Google$0.01$1.0K131K
152gpt-5.2OpenAI$0.01$1.0K400K
153gpt-5.2-chat-latestOpenAI$0.01$1.0K
154gpt-5.3-chat-latestOpenAI$0.01$1.0K
155gpt-5.3-codexOpenAI$0.01$1.0K400K
156gpt-4.1-2025-04-14OpenAI$0.01$1.1K
157MoonshotAI Kimi LatestVia OpenRouterMoonshot AI (Kimi)$0.01$1.1K1M
158gpt-5.4OpenAI$0.01$1.2K1.1M
159Claude Sonnet 4Anthropic$0.01$1.2K1M
160Claude Sonnet 4.5Anthropic$0.01$1.2K1M
161Claude Sonnet 4.6Anthropic$0.01$1.2K1M
162Kimi K3Via OpenRouterMoonshot AI (Kimi)$0.01$1.2K1M
163Gemini 2.5 Pro Preview TTSGoogle$0.01$1.4K
164Gemini 3.1 Flash TTS PreviewGoogle$0.01$1.4K
165gpt-4o-2024-08-06OpenAI$0.01$1.4K128K
166o4-mini-2025-04-16OpenAI$0.01$1.4K
167gpt-4o-2024-05-13OpenAI$0.02$1.6K128K
168Gemini 3.5 Live TranslateGoogle$0.02$1.8K
169Claude Opus 4.5Anthropic$0.02$2.0K200K
170Claude Opus 4.6Anthropic$0.02$2.0K1M
171Claude Opus 4.7Anthropic$0.02$2.0K1M
172Claude Opus 4.8Anthropic$0.02$2.0K1M
173Claude Opus 5Anthropic$0.02$2.0K1M
174chat-latestOpenAI$0.02$2.3K
175gpt-5.5OpenAI$0.02$2.3K1.1M
176gpt-5.6-solOpenAI$0.02$2.3K1.1M
177gpt-image-2OpenAI$0.03$2.7K
178chatgpt-image-latestOpenAI$0.03$2.9K
179gpt-image-1.5OpenAI$0.03$2.9K
180gpt-4-turbo-2024-04-09OpenAI$0.03$3.3K
181gpt-image-1OpenAI$0.04$3.6K
182Claude Fable 5Anthropic$0.04$4.1K1M
183Claude Mythos 5Anthropic$0.04$4.1K
184o1OpenAI$0.06$5.5K200K
185gpt-5.5-cyberOpenAI$0.06$5.9K
186gpt-5.6-cyberOpenAI$0.06$5.9K
187Claude Opus 4Anthropic$0.06$6.1K200K
188Claude Opus 4.1Anthropic$0.06$6.1K200K
189o3-proOpenAI$0.08$7.8K200K
190gpt-4-0613OpenAI$0.08$8.1K
191gpt-5-proOpenAI$0.09$9.4K400K
192gpt-5.2-proOpenAI$0.13$13.2K400K
193gpt-5.4-proOpenAI$0.15$15.3K1.1M
194gpt-5.5-proOpenAI$0.15$15.3K1.1M
195o1-proOpenAI$0.58$58.5K200K

18 models excluded because they cannot serve this workload — mostly embedding, moderation and OCR endpoints that publish no output price because they do not generate text. Counting that missing price as zero would rank them as the cheapest way to run a chat.

Estimates use standard-tier list prices. They exclude batch discounts, committed-use agreements and free allowances, and assume every request is the same shape. Treat the ranking as sound and the absolute figures as approximate.

Questions about these estimates

Why does the cheapest model change when I change the workload?
Because output tokens usually cost three to five times what input tokens cost. A model with cheap input and expensive output wins for retrieval-style work with short answers and loses badly for anything that generates long text. Ranking on list price alone hides that.
What does the cached input percentage do?
Most providers bill a reduced rate — often around a tenth of the normal input price — when a prompt repeats a prefix they have already processed. If you send the same system prompt or document on every call, that share is billed at the cheaper rate. It is usually the single largest saving available.
Why are some rows greyed out?
Their context window is smaller than the tokens you entered, so they cannot hold your prompt. They are shown rather than hidden, because a model that cannot do the job is not a cheaper option.
How accurate are these numbers?
They use standard-tier list prices updated daily, and assume every request is the same shape. They exclude batch discounts, committed-use agreements and free allowances. Treat the ranking as sound and the absolute totals as an estimate.

Compare all 216 models by list price →