Skip to content
LLM Toolkit

LLM API Pricing Comparison

Every major model, priced side by side. 22 models from OpenAI, Anthropic and Google — input, cached input and output rates per 1M tokens, plus context windows. Prices verified against official provider pages as of 2026-08-26; click any column header to sort.

Cheapest input

GPT-5 nano

$0.050 /1M in

Cheapest output

GPT-5 nano

$0.40 /1M out

Largest context

Claude Fable 5

1M tokens

All 22 models, sorted by input price

Model Context Input /1M Cached /1M Output /1M Verified
GPT-5 nano OpenAI 400K $0.050 $0.005 $0.40 2026-08-26
GPT-4o mini OpenAI 128K $0.15 $0.075 $0.60 2026-08-26
GPT-5.6 Luna OpenAI 400K $0.20 $0.020 $1.20 2026-08-26
GPT-5.4 nano OpenAI 400K $0.20 $0.020 $1.25 2026-08-26
GPT-5 mini OpenAI 400K $0.25 $0.025 $2.00 2026-08-26
Gemini 3.5 Flash-Lite Google 1M $0.30 $0.030 $2.50 2026-08-26
GPT-5.4 mini OpenAI 400K $0.75 $0.075 $4.50 2026-08-26
Gemini 3.7 Flash Google 1M $0.75 $0.075 $3.75 2026-08-26
Claude Haiku 4.5 Anthropic 200K $1.00 $0.10 $5.00 2026-08-26
o4-mini OpenAI 200K $1.10 $0.28 $4.40 2026-08-26
GPT-5 OpenAI 400K $1.25 $0.13 $10.00 2026-08-26
Gemini 3.5 Flash Google 1M $1.50 $0.15 $9.00 2026-08-26
GPT-5.2 OpenAI 400K $1.75 $0.17 $14.00 2026-08-26
GPT-5.6 Terra OpenAI 400K $2.00 $0.20 $12.00 2026-08-26
o3 OpenAI 200K $2.00 $0.50 $8.00 2026-08-26
Claude Sonnet 5 Anthropic 1M $2.00 $0.20 $10.00 2026-08-26
GPT-5.4 OpenAI 400K $2.50 $0.25 $15.00 2026-08-26
GPT-4o OpenAI 128K $2.50 $1.25 $10.00 2026-08-26
GPT-5.6 Sol OpenAI 400K $5.00 $0.50 $30.00 2026-08-26
GPT-5.5 OpenAI 400K $5.00 $0.50 $30.00 2026-08-26
Claude Opus 5 Anthropic 1M $5.00 $0.50 $25.00 2026-08-26
Claude Fable 5 Anthropic 1M $10.00 $1.00 $50.00 2026-08-26

Prices are per 1 million tokens in USD, from official provider pricing pages. A green dot means the row was checked against the provider's own page on the date shown; an amber dot means the figure comes from a secondary source and awaits verification. Llama and other open-weight models are served at provider-dependent rates and are excluded until we can verify them.

How we keep this table honest

Most pricing trackers quietly rot: providers change rates, launch tiers, retire models — and the table keeps showing last year's numbers. We do three things differently. Every row carries its own verification date, checked against the provider's official pricing page. Rows we haven't verified are labeled as such instead of pretending confidence. And the whole table gets re-checked on the first of every month — including after major model launches, when prices tend to move.

Want to sanity-check your own numbers before committing to a provider? Count your real token usage with the AI Token Counter — it runs the exact OpenAI tokenizer in your browser, so the counts match your bill.

Frequently asked questions

What is the cheapest LLM API right now?

As of August 2026, the cheapest verified production-grade API is GPT-5 nano at $0.050 per 1M input tokens and $0.40 per 1M output tokens. If your workload is output-heavy, GPT-5 nano is the cheapest per output token at $0.40 per 1M. Prices are verified against official provider pricing pages.

How much does GPT-5 cost per million tokens?

GPT-5 costs $1.25 per 1M input tokens and $10.00 per 1M output tokens, with cached input at $0.125. The smaller tiers are far cheaper: GPT-5 mini runs $0.25 in / $2.00 out, and GPT-5 nano runs $0.05 in / $0.40 out — 25× cheaper than the flagship on input.

How much does Claude cost per million tokens?

Anthropic’s current lineup: Claude Fable 5 at $10.00 in / $50.00 out, Claude Opus 5 at $5.00 in / $25.00 out, Claude Sonnet 5 at $2.00 in / $10.00 out (the scheduled increase to $3/$15 was cancelled), and Claude Haiku 4.5 at $1.00 in / $5.00 out. Claude 5-series models ship a 1M token context window — Haiku 4.5 stays at 200K — and cached input costs 10% of the base input rate.

What does cached input pricing mean?

Prompt caching lets you reuse the input you’ve already sent — a long system prompt, a document, or conversation history — at a steep discount, typically 10% of the normal input price. If your requests repeat most of their context (agents, RAG, multi-turn chat), caching routinely cuts input costs by 60–90%. The cached column in the table shows each model’s discounted rate.

Which LLM has the largest context window?

Claude Fable 5 leads with a 1M-token context window. Google’s Gemini family and Anthropic’s Claude 5 series ship 1M tokens across their lineups, while OpenAI’s GPT-5 family offers 400K. More context costs more per request in practice, because you pay for every input token you send.

How should I pick a model based on price?

Match the model tier to the task difficulty, not the other way around. Route high-volume, simple tasks (classification, extraction, routing) to nano-tier models; save flagship models for hard reasoning. Then check your input/output ratio: agent workloads that generate long outputs should compare output prices first — that’s usually 80% of the bill. Our API Cost Calculator (shipping next) computes this for your exact volumes.

Next in the toolkit

The API Cost Calculator — turn these per-token rates into a real monthly bill for your request volumes — plus Context Window Comparison and the Token ↔ Words Converter. See all tools