LLM API Pricing Comparison
Every major model, priced side by side. 22 models from OpenAI, Anthropic and Google — input, cached input and output rates per 1M tokens, plus context windows. Prices verified against official provider pages as of 2026-08-26; click any column header to sort.
Cheapest input
GPT-5 nano
$0.050 /1M in
Cheapest output
GPT-5 nano
$0.40 /1M out
Largest context
Claude Fable 5
1M tokens
All 22 models, sorted by input price
| Model | Context | Input /1M | Cached /1M | Output /1M | Verified |
|---|---|---|---|---|---|
| GPT-5 nano OpenAI | 400K | $0.050 | $0.005 | $0.40 | 2026-08-26 |
| GPT-4o mini OpenAI | 128K | $0.15 | $0.075 | $0.60 | 2026-08-26 |
| GPT-5.6 Luna OpenAI | 400K | $0.20 | $0.020 | $1.20 | 2026-08-26 |
| GPT-5.4 nano OpenAI | 400K | $0.20 | $0.020 | $1.25 | 2026-08-26 |
| GPT-5 mini OpenAI | 400K | $0.25 | $0.025 | $2.00 | 2026-08-26 |
| Gemini 3.5 Flash-Lite Google | 1M | $0.30 | $0.030 | $2.50 | 2026-08-26 |
| GPT-5.4 mini OpenAI | 400K | $0.75 | $0.075 | $4.50 | 2026-08-26 |
| Gemini 3.7 Flash Google | 1M | $0.75 | $0.075 | $3.75 | 2026-08-26 |
| Claude Haiku 4.5 Anthropic | 200K | $1.00 | $0.10 | $5.00 | 2026-08-26 |
| o4-mini OpenAI | 200K | $1.10 | $0.28 | $4.40 | 2026-08-26 |
| GPT-5 OpenAI | 400K | $1.25 | $0.13 | $10.00 | 2026-08-26 |
| Gemini 3.5 Flash Google | 1M | $1.50 | $0.15 | $9.00 | 2026-08-26 |
| GPT-5.2 OpenAI | 400K | $1.75 | $0.17 | $14.00 | 2026-08-26 |
| GPT-5.6 Terra OpenAI | 400K | $2.00 | $0.20 | $12.00 | 2026-08-26 |
| o3 OpenAI | 200K | $2.00 | $0.50 | $8.00 | 2026-08-26 |
| Claude Sonnet 5 Anthropic | 1M | $2.00 | $0.20 | $10.00 | 2026-08-26 |
| GPT-5.4 OpenAI | 400K | $2.50 | $0.25 | $15.00 | 2026-08-26 |
| GPT-4o OpenAI | 128K | $2.50 | $1.25 | $10.00 | 2026-08-26 |
| GPT-5.6 Sol OpenAI | 400K | $5.00 | $0.50 | $30.00 | 2026-08-26 |
| GPT-5.5 OpenAI | 400K | $5.00 | $0.50 | $30.00 | 2026-08-26 |
| Claude Opus 5 Anthropic | 1M | $5.00 | $0.50 | $25.00 | 2026-08-26 |
| Claude Fable 5 Anthropic | 1M | $10.00 | $1.00 | $50.00 | 2026-08-26 |
Prices are per 1 million tokens in USD, from official provider pricing pages. A green dot means the row was checked against the provider's own page on the date shown; an amber dot means the figure comes from a secondary source and awaits verification. Llama and other open-weight models are served at provider-dependent rates and are excluded until we can verify them.
How we keep this table honest
Most pricing trackers quietly rot: providers change rates, launch tiers, retire models — and the table keeps showing last year's numbers. We do three things differently. Every row carries its own verification date, checked against the provider's official pricing page. Rows we haven't verified are labeled as such instead of pretending confidence. And the whole table gets re-checked on the first of every month — including after major model launches, when prices tend to move.
Want to sanity-check your own numbers before committing to a provider? Count your real token usage with the AI Token Counter — it runs the exact OpenAI tokenizer in your browser, so the counts match your bill.
Frequently asked questions
What is the cheapest LLM API right now?
As of August 2026, the cheapest verified production-grade API is GPT-5 nano at $0.050 per 1M input tokens and $0.40 per 1M output tokens. If your workload is output-heavy, GPT-5 nano is the cheapest per output token at $0.40 per 1M. Prices are verified against official provider pricing pages.
How much does GPT-5 cost per million tokens?
GPT-5 costs $1.25 per 1M input tokens and $10.00 per 1M output tokens, with cached input at $0.125. The smaller tiers are far cheaper: GPT-5 mini runs $0.25 in / $2.00 out, and GPT-5 nano runs $0.05 in / $0.40 out — 25× cheaper than the flagship on input.
How much does Claude cost per million tokens?
Anthropic’s current lineup: Claude Fable 5 at $10.00 in / $50.00 out, Claude Opus 5 at $5.00 in / $25.00 out, Claude Sonnet 5 at $2.00 in / $10.00 out (the scheduled increase to $3/$15 was cancelled), and Claude Haiku 4.5 at $1.00 in / $5.00 out. Claude 5-series models ship a 1M token context window — Haiku 4.5 stays at 200K — and cached input costs 10% of the base input rate.
What does cached input pricing mean?
Prompt caching lets you reuse the input you’ve already sent — a long system prompt, a document, or conversation history — at a steep discount, typically 10% of the normal input price. If your requests repeat most of their context (agents, RAG, multi-turn chat), caching routinely cuts input costs by 60–90%. The cached column in the table shows each model’s discounted rate.
Which LLM has the largest context window?
Claude Fable 5 leads with a 1M-token context window. Google’s Gemini family and Anthropic’s Claude 5 series ship 1M tokens across their lineups, while OpenAI’s GPT-5 family offers 400K. More context costs more per request in practice, because you pay for every input token you send.
How should I pick a model based on price?
Match the model tier to the task difficulty, not the other way around. Route high-volume, simple tasks (classification, extraction, routing) to nano-tier models; save flagship models for hard reasoning. Then check your input/output ratio: agent workloads that generate long outputs should compare output prices first — that’s usually 80% of the bill. Our API Cost Calculator (shipping next) computes this for your exact volumes.
Next in the toolkit
The API Cost Calculator — turn these per-token rates into a real monthly bill for your request volumes — plus Context Window Comparison and the Token ↔ Words Converter. See all tools