LLM Token Counter & API Cost Calculator

Paste a prompt. Get an exact token count for OpenAI models and an honest estimate for Claude, Gemini, and DeepSeek — then see what that same text would cost on 14 models, side by side.

…tokens exact
0characters
0words

What this text would cost, per model

Sorted cheapest first. Assumes 0 input tokens and 500 output tokens. Counts for non-OpenAI rows are estimates.

ModelInput / 1MOutput / 1MEst. costWith cache hits
Gemini 2.5 Flash-Lite · Google$0.1$0.4$0.0002—
GPT-6 Luna · OpenAI$0.1$0.5$0.00025$0.00025
DeepSeek-V4.1 Flash · DeepSeek$0.15$0.6$0.0003$0.0003
DeepSeek-V4 Pro · DeepSeek$0.66$1.98$0.00099—
Gemini 3.8 Flash · Google$0.75$3.75$0.001875—
Claude Haiku 4.5 · Anthropic$1$5$0.0025$0.0025
GPT-6 Sol · OpenAI$2$10$0.005$0.005
Claude Sonnet 5 · Anthropic$2$10$0.005$0.005
Claude Sonnet 5.5 · Anthropic$2$10$0.005—
Gemini 3.1 Pro Preview · Google$2$12$0.006—
Claude Opus 5.5 · Anthropic$4$20$0.01$0.01
Claude Opus 5 · Anthropic$5$25$0.0125$0.0125
GPT-6 Astra · OpenAI$10$50$0.025—
Claude Fable 5.1 · Anthropic$10$50$0.025$0.025

“With cache hits” assumes all input tokens hit the prompt cache, where the provider publishes a cached-input rate. DeepSeek rows use off-peak rates; peak hours cost 2×.

Prices change frequently — last verified 2026-09-30. Rates below are USD per 1M tokens from public pricing pages. Always confirm on the provider's official pricing page before budgeting.

How it works

  • OpenAI: exact counts using the real o200k_base (GPT-6 / GPT-5) and cl100k_base (GPT-4 and older) encodings, run in your browser.
  • Claude / Gemini / DeepSeek: character-based heuristic estimates, always labeled as estimates — because these providers haven't published a client-side tokenizer.
  • Costs: your token count × public per-million-token rates, with an expected-output length you control and a cache-hit column where providers publish one.

Counting a single model? Use the dedicated GPT token counter, Claude token counter, Gemini token counter, or DeepSeek token counter. For rates alone, see the LLM API pricing comparison or the OpenAI pricing calculator.

Frequently asked questions

Is the token count exact?

For OpenAI models it is: we run the same o200k_base and cl100k_base encodings OpenAI uses, entirely in your browser. For Claude, Gemini, and DeepSeek we show a clearly-labeled heuristic estimate, because those providers have not published standalone tokenizers you can run client-side.

Why do different models count tokens differently?

Each provider trains its own tokenizer, so the same English sentence can split into a different number of tokens on GPT-6, Claude, and Gemini. That is also why per-token prices are not directly comparable — a cheaper per-token price can still cost more if the model uses more tokens for the same text.

How is the cost table calculated?

Cost = (your input tokens ÷ 1,000,000 × input price) + (expected output tokens ÷ 1,000,000 × output price). The “with cache hits” column assumes all input tokens hit the prompt cache, using the provider's published cached-input rate where one exists.

Do my prompts leave my browser?

No. Counting and cost math all run locally in JavaScript. Nothing is uploaded, logged, or sent anywhere — the only exception is the optional “exact count via official API” panels, which call the provider's API directly with a key you supply.

How often are prices updated?

Prices are checked against public pricing pages and updated regularly; the last verification date is shown above the table. Model pricing changes often, so confirm on the provider's official pricing page before making budget decisions.

What is prompt caching and why does it matter?

If you resend the same long prefix (system prompt, documents) across requests, providers can serve it from cache at a steep discount — typically 90% off input price, and 98% off on DeepSeek. For agents that reread the same context all day, cache hit rate often affects the bill more than the headline price.