Cheapest input
- 1GLM-4.7-FlashXZ.ai$0.070per 1M tokens
- 2Qwen3.5 FlashAlibaba$0.100per 1M tokens
- 3GPT-4o MiniOpenAI$0.150per 1M tokens
- 4GLM-4.5-AirZ.ai$0.200per 1M tokens
- 5GPT-5.6 LunaOpenAI$0.200per 1M tokens
Every rate here was read from the provider's own pricing page on August 22, 2026. Compare input, cached and output prices across 56 models from 11 providers — then enter your own token volumes to see what a single request would cost on each one.
Click a column heading to sort.
| per 1M | per 1M | per 1M | Counting | per request | ||
|---|---|---|---|---|---|---|
| GPT-5.6 Sol1OpenAI | $4.00 | $0.400 | $20.00 | 1M | exact | $0.060 |
| GPT-5.6 Terra1OpenAI | $2.00 | $0.200 | $12.00 | 1M | exact | $0.032 |
| GPT-5.6 Luna1OpenAI | $0.200 | $0.020 | $1.20 | 1M | exact | $0.00320 |
| GPT-5.51OpenAI | $5.00 | $0.500 | $30.00 | 1M | exact | $0.080 |
| GPT-5.41OpenAI | $2.50 | $0.250 | $15.00 | 1M | exact | $0.040 |
| GPT-5.4 mini1OpenAI | $0.750 | $0.075 | $4.50 | 400K | exact | $0.012 |
| GPT-5.4 nano1OpenAI | $0.200 | $0.020 | $1.25 | 400K | exact | $0.00325 |
| GPT-5.11OpenAI | $1.25 | $0.125 | $10.00 | 400K | exact | $0.022 |
| GPT-4.11OpenAI | $2.00 | $0.500 | $8.00 | 1M | exact | $0.028 |
| GPT-4o1OpenAI | $2.50 | $1.25 | $10.00 | 128K | exact | $0.035 |
| GPT-4o Mini1OpenAI | $0.150 | $0.075 | $0.600 | 128K | exact | $0.00210 |
| Claude Fable 52Anthropic | $10.00 | $1.00 | $50.00 | 1M | est. | $0.150 |
| Claude Opus 52Anthropic | $5.00 | $0.500 | $25.00 | 1M | est. | $0.075 |
| Claude Opus 4.82Anthropic | $5.00 | $0.500 | $25.00 | 1M | est. | $0.075 |
| Claude Opus 4.72Anthropic | $5.00 | $0.500 | $25.00 | 1M | est. | $0.075 |
| Claude Opus 4.62Anthropic | $5.00 | $0.500 | $25.00 | 1M | est. | $0.075 |
| Claude Opus 4.52Anthropic | $5.00 | $0.500 | $25.00 | 200K | est. | $0.075 |
| Claude Sonnet 52Anthropic | $2.00 | $0.200 | $10.00 | 1M | est. | $0.030 |
| Claude Sonnet 4.62Anthropic | $3.00 | $0.300 | $15.00 | 1M | est. | $0.045 |
| Claude Sonnet 4.52Anthropic | $3.00 | $0.300 | $15.00 | 200K | est. | $0.045 |
| Claude Haiku 4.52Anthropic | $1.00 | $0.100 | $5.00 | 200K | est. | $0.015 |
| Gemini 3.1 Pro3Google | $2.00 | $0.200 | $12.00 | 1M | est. | $0.032 |
| Gemini 3.7 FlashGoogle | $0.750 | $0.075 | $3.75 | 1M | est. | $0.011 |
| Gemini 3.5 FlashGoogle | $1.50 | $0.150 | $9.00 | 1M | est. | $0.024 |
| Gemini 3.5 Flash-LiteGoogle | $0.300 | $0.030 | $2.50 | 1M | est. | $0.00550 |
| Gemini 2.5 Pro3Google | $1.25 | $0.125 | $10.00 | 1M | est. | $0.022 |
| Grok 4.64xAI | $2.00 | $0.500 | $6.00 | 500K | est. | $0.026 |
| Grok 4.54xAI | $2.00 | $0.300 | $6.00 | 500K | est. | $0.026 |
| Grok 4.34xAI | $1.25 | $0.200 | $2.50 | 1M | est. | $0.015 |
| Grok Build 0.14xAI | $1.00 | $0.200 | $2.00 | 256K | est. | $0.012 |
| DeepSeek V4 Pro5DeepSeek | $1.32 | $0.044 | $3.96 | 1M | est. | $0.017 |
| DeepSeek V4 Flash5DeepSeek | $0.440 | $0.014 | $1.32 | 1M | est. | $0.00572 |
| LLaMA 3.3 70B6Meta (via API) | $1.04 | — | $1.04 | 128K | est. | $0.011 |
| Qwen3.7 Max1617Alibaba | $2.50 | $0.250 | $7.50 | 1M | est. | $0.033 |
| Sonar7918Perplexity | $1.00 | — | $1.00 | — | est. | $0.011 |
| Sonar Pro7918Perplexity | $3.00 | — | $15.00 | — | est. | $0.045 |
| Sonar Reasoning Pro918Perplexity | $2.00 | — | $8.00 | — | est. | $0.028 |
| Sonar Deep Research8918Perplexity | $2.00 | — | $8.00 | — | est. | $0.028 |
| Kimi K310Moonshot AI | $3.00 | $0.300 | $15.00 | 1M | est. | $0.045 |
| Kimi K2.7 Code1011Moonshot AI | $0.950 | $0.190 | $4.00 | 262K | est. | $0.013 |
| Kimi K2.610Moonshot AI | $0.950 | $0.160 | $4.00 | 262K | est. | $0.013 |
| MiniMax M31213MiniMax | $0.300 | $0.060 | $1.20 | 1M | est. | $0.00420 |
| MiniMax M2.714MiniMax | $0.300 | $0.060 | $1.20 | 205K | est. | $0.00420 |
| MiniMax M2.5MiniMax | $0.300 | $0.030 | $1.20 | 205K | est. | $0.00420 |
| GLM-5.315Z.ai | $1.40 | $0.280 | $4.40 | 1M | est. | $0.018 |
| GLM-5.215Z.ai | $1.40 | $0.280 | $4.40 | 1M | est. | $0.018 |
| GLM-515Z.ai | $1.00 | $0.200 | $3.20 | 200K | est. | $0.013 |
| GLM-5-Turbo15Z.ai | $1.20 | $0.240 | $4.00 | 200K | est. | $0.016 |
| GLM-4.715Z.ai | $0.600 | $0.120 | $2.20 | 200K | est. | $0.00820 |
| GLM-4.7-FlashX15Z.ai | $0.070 | $0.014 | $0.400 | 200K | est. | $0.00110 |
| GLM-4.5-Air15Z.ai | $0.200 | $0.040 | $1.10 | 128K | est. | $0.00310 |
| GLM-4.5-X1518Z.ai | $2.20 | $0.440 | $8.90 | — | est. | $0.031 |
| Qwen3.7 Plus1617Alibaba | $0.400 | $0.040 | $1.60 | 1M | est. | $0.00560 |
| Qwen3.6 Flash1617Alibaba | $0.250 | $0.025 | $1.50 | 1M | est. | $0.00400 |
| Qwen3.5 Flash17Alibaba | $0.100 | $0.010 | $0.400 | 1M | est. | $0.00140 |
| Qwen3 Coder Plus1617Alibaba | $1.00 | $0.100 | $5.00 | 1M | est. | $0.015 |
| No models match that search. | ||||||
All rates are in USD per 1M tokens.
Checked against each provider’s own pricing page on August 22, 2026.
Providers change prices without notice, and several bill extra beyond the token rate — confirm against the provider’s page before you budget.
Output rates in this table run from $0.400 to $50.00 per 1M tokens — a 125× spread, which is a bigger lever than any prompt tuning.
11 providers, each linked to the pricing page these figures were read from.
On input, GLM-4.7-FlashX at $0.070 per 1M tokens. On output, GLM-4.7-FlashX at $0.400 per 1M. Which one is cheapest for you depends on the shape of your traffic: a 10,000-token prompt with a 1,000-token answer costs $0.00110 on GLM-4.7-FlashX and $0.032 on GPT-5.6 Terra. Set the two token fields above to your own volumes and the table re-sorts around them.
Weight input and output by how you actually use the model. Providers charge three to five times more for what the model writes than for what you send — GPT-5.6 Terra is $2.00 in and $12.00 out — so a summariser that reads a lot and writes a little ranks the models completely differently from a code generator that does the opposite. Enter both volumes above rather than comparing input rates alone, and check the notes: several providers re-price the entire request once a prompt crosses an input threshold.
When consecutive requests share the same prefix — a system prompt, a document, a long few-shot block — providers can serve that prefix from cache and charge less for it. The cached column is that cache-hit read rate, usually 10% to 20% of the standard input rate: Claude Sonnet 5 reads cached input at $0.200 against $2.00 standard. Writing to the cache is not always free, and a dash means the provider publishes no cached rate at all.
Every figure was read from the provider's own pricing page on August 22, 2026, and each provider card links back to the page it came from. Nothing here is scraped live, so treat it as a snapshot: providers change rates without notice, and the ones with promotional or off-peak pricing (MiniMax, DeepSeek) move most often.
Because the provider does not publish that figure, and guessing would be worse than leaving it blank. Perplexity publishes no context window for the Sonar models, Z.ai publishes no per-variant context length for GLM-4.5-X, and Meta’s hosted LLaMA row has no cached rate. A dash always means "not published", never "zero".
11 of the 56 rows — the OpenAI models, because OpenAI publishes its tokenizer and the calculator loads the real BPE vocabulary (o200k_base) in your browser. Everything else is marked "est." and counted with a characters-per-token approximation, which lands within roughly 10–15% on English prose but drifts on code, JSON and non-Latin scripts. The prices are first-party for every row; only the token counts differ in confidence.
No — the table shows token rates only, and a few providers bill on top of them. Sonar adds $5 to $12 per 1,000 requests for search depending on context size, Sonar Deep Research adds citation-token, reasoning-token and per-search charges, and hosted open-weight models carry whatever the host charges rather than a first-party rate. The notes column flags every row where the token rate is not the whole bill.
Take 200,000 input tokens — a full novel, or a few hundred pages of a contract. That single prompt costs $0.400 on GPT-5.6 Terra, $0.400 on Claude Sonnet 5, $0.600 on Kimi K3 and $0.020 on Qwen3.5 Flash. Watch the tier notes at that size: on Grok 4.6 a prompt of 200K tokens bills the whole request at double the listed rates.
A price per million tokens only means something once you know how many tokens your text is. The token calculator counts as you type and prices the result across models; the file counter does the same for PDFs and Word documents.