Fresh paint, wet floors — we're rebuilding CostOfToken into a cross-provider price comparison. Some corners are still under construction; the full version is coming up soon.

CostOfToken

Updated daily

Updated 16 hours ago · No signup · Free & open · LLM API pricing normalized to USD per 1M tokens

Average across 473 selected models

$0.292
Avg input /1M
$1.24
Avg output /1M
$0.766
Avg blended /1M

Showing 3 popular of 473 models · select one for details, or tick up to 3 to compare

Same model, different bill

Popular models priced by every seller we track, side by side — USD per 1M tokens, standard tier. The cheapest door is highlighted; click any price for that seller's page. Free and promo routes live on /free and /discounts.

Input and output price per 1M tokens for popular models across sellers
ModelFirst-partyin / out per 1MOpenRouterin / out per 1MTogether AIin / out per 1MDeepInfrain / out per 1M
gpt-5.6-sol$4.00 / $20.00$2.00 / $10.00cheapest——
Claude Opus 5$5.00 / $25.00cheapest$5.00 / $25.00—$5.00 / $25.00
Gemini 3.1 Pro Preview$2.00 / $12.00cheapest$2.00 / $12.00——
grok-4.5$2.00 / $6.00cheapest$2.00 / $6.00——
Claude Fable 5$10.00 / $50.00cheapest$10.00 / $50.00—$10.00 / $50.00
Claude Sonnet 5$2.00 / $10.00cheapest$2.00 / $10.00—$3.00 / $15.00
gpt-5.6-luna$0.200 / $1.20cheapest$0.200 / $1.20——
Gemini 3.6 Flash$0.750 / $3.75cheapest$0.750 / $3.75——
DeepSeek V4 Pro 0423$0.789 / $1.58cheapest$0.789 / $1.58——
GLM-5.2$1.40 / $4.40$0.650 / $2.04cheapest$1.40 / $4.40—

LLM pricing questions

Which LLM API is cheapest?
It depends on the mix of input and output tokens, since output is typically 3-5x the price of input. Sort the table by blended cost to rank on a simple mean of the two, or by input alone if your workload is prompt-heavy. Several Zhipu GLM Flash models are genuinely free.
What does "per 1M tokens" mean?
Providers bill per token, and a token is roughly 0.75 of an English word. Quoting per 1,000,000 tokens is the industry convention because per-token prices run to eight decimal places. Every price here is normalized to USD per 1M tokens so models are directly comparable.
What is cached input pricing?
Most providers charge a reduced rate when a prompt repeats a prefix they have already processed — often 10% of the normal input price. If you send the same system prompt or document on every call, cached input is usually the single largest saving available.
How current are these prices?
They are re-read from each provider every day. A price is only recorded when it actually changes, so the history shows real movements rather than daily noise, and every row shows the date it was last confirmed.
Is there an API for this pricing data?
Yes. GET /api/v1/prices returns the whole table as JSON, free and without signup, rate limited to 60 requests per hour per IP. /llms-full.txt serves the same data as markdown for LLM ingestion.