LLM API Pricing Comparison

Every major model's per-token price — input, output, and cached input per 1M tokens — in one sortable table. Filter by provider, click any column to sort.

Prices change frequently — last verified 2026-09-30. Rates below are USD per 1M tokens from public pricing pages. Always confirm on the provider's official pricing page before budgeting.

GPT-6 LunaCached input reads get a 90% discount.OpenAI$0.1$0.5$0.012026-09-23
Gemini 2.5 Flash-LiteGoogle$0.1$0.4—2026-09-25
DeepSeek-V4.1 FlashOff-peak rates shown. Peak-hour traffic costs 2x. Cached input billed at 2% of the input rate.DeepSeek$0.15$0.6$0.0032026-09-27
DeepSeek-V4 ProOff-peak rates shown. Peak-hour traffic costs 2x.DeepSeek$0.66$1.98—2026-09-27
Gemini 3.8 FlashIntroductory pricing through December 31, 2026.Google$0.75$3.75—2026-09-29
Claude Haiku 4.5Anthropic$1$5$0.12026-09-24
GPT-6 SolCached input reads get a 90% discount.OpenAI$2$10$0.22026-09-23
Claude Sonnet 5Anthropic$2$10$0.22026-09-24
Claude Sonnet 5.5Anthropic$2$10—2026-09-29
Gemini 3.1 Pro PreviewInput price doubles above 200K input tokens.Google$2$12—2026-09-29
Claude Opus 5.5Anthropic$4$20$0.22026-09-24
Claude Opus 5Anthropic$5$25$0.52026-09-24
GPT-6 AstraStandard tier, prompts up to 272K input tokens; higher long-context rates apply beyond that.OpenAI$10$50—2026-09-22
Claude Fable 5.1Cache reads billed at 2.5% of the input price on Fable models.Anthropic$10$50$0.252026-09-24

Click a column header to sort. “—” means the provider does not publish a cached-input rate. DeepSeek rows show off-peak rates; peak hours cost 2×.

Open the token cost calculator → Paste your prompt and see what it costs on each of these models.

How to read this table

  • Input / Output per 1M are what you pay per million tokens sent and generated. Output tokens almost always cost more — often 4–5× the input rate.
  • Cached input is the discounted rate for repeated prompt prefixes. If your workload reuses long system prompts, sort by this column instead.
  • Per-token price ≠ per-task price. A model at $0.10 per 1M that needs 3× the tokens can lose to a $2 model that answers in fewer tokens — especially for reasoning models that “think” before answering.

Counting tokens for one provider? Try the GPT token counter, Claude token counter, Gemini token counter, or DeepSeek token counter.

Frequently asked questions

Which is the cheapest LLM API right now?

On headline per-token rates, DeepSeek's off-peak pricing is usually the lowest, followed by Google's Flash-Lite tier. But headline rates mislead: tokenizers differ, so the same task can use very different token counts per model. Use the token counter to measure your actual prompt, then multiply — the table above sorts by input price so you can see the ranking at a glance.

Why do some models cost 10x more than others?

You are paying for capability and scale. Frontier reasoning models run far more compute per token than small flash models. For classification, extraction, or summarization, a cheap model is often indistinguishable from a flagship — the expensive models earn their price on hard reasoning, coding, and agentic tasks.

What is prompt caching and how much does it save?

When you resend the same long prefix (system prompt, documents) across requests, providers serve it from cache at a steep discount — typically 90% off the input rate, and 98% off on DeepSeek. For agents that reread the same context all day, cache hit rate usually affects the bill more than the headline price.

Are these the official prices?

Rates are taken from public pricing pages and reputable roundups, and each row shows the date it was last checked. Prices change often — some, like Gemini's introductory tiers, are explicitly time-limited — so confirm on the provider's official pricing page before making budget decisions.

How do I turn these rates into a real cost estimate?

Count the tokens in your actual prompt with the token counter, plug the input/output counts into the cost table, and you get a per-request figure for every model. For OpenAI-only math, the dedicated OpenAI pricing calculator does it in one step.