Tool · prices checked 2026-10-01

LLM API cost calculator

Put in your traffic and typical prompt size, and see what each current Claude, GPT and Gemini model would cost per month — before and after prompt caching and batch discounts.

ModelPer requestPer monthvs cheapest

How the calculator works

For each model it computes input tokens × input price + output tokens × output price per request, using prices per million tokens from each provider's official page. If you set a cache share, that part of the input is billed at the model's cached-input rate instead. Batch mode halves the result, which matches the 50% batch discount Anthropic, OpenAI and Google all publish. A month is 30 days.

It deliberately leaves out a few things that are real but small or hard to predict: one-time cache-write premiums, Google's hourly cache storage fee, long-context surcharges, tool-call tokens, regional uplifts and taxes. For a typical chat or RAG workload they move the total by a few percent; for very long prompts check the long-context rows on the pricing table.

Getting realistic token numbers

The most common mistake is guessing token counts. A token is roughly ¾ of an English word, but non-English text, code and JSON use more. Two practical ways to get real numbers:

  • Read the usage field. Every API response reports input and output tokens. Log them for a day of real traffic and use the median, not the maximum.
  • Count before you send. Anthropic and Google both offer free token-counting endpoints, and OpenAI documents its tokenizer. Count your system prompt once — it is usually the biggest single chunk and the best candidate for caching.

Also note that Anthropic says its Claude 4.7-and-later tokenizer produces roughly 30% more tokens for the same text than earlier Claude models, so token counts are not portable between providers or even model generations.

Three levers that cut the bill

  1. Cache the stable prefix. Put the system prompt, tool definitions and reference documents first and keep them byte-identical between calls. With 80% of input cached, input cost on most models drops by about 70%.
  2. Batch anything that can wait. Reports, nightly enrichment, evaluation runs: 50% off with no code change beyond the endpoint.
  3. Route by difficulty. Send every request to a small model first and escalate only failures. In the table above the gap between the cheapest and the most expensive model is often 50–100×.

FAQ

Are these prices current?

They were checked against the official pages on 2026-10-01. Our build fails if any of the 20 prices stops matching the source page, so the site is rebuilt with corrected numbers rather than silently going stale.

Why does the cheapest model differ from what I expected?

Because the ratio of input to output matters. Models with cheap input but expensive output win on long-document tasks and lose on long-answer tasks. Try changing the output tokens and watch the ranking move.

Does this include free tiers?

No. Free tiers have rate limits and, for some providers, allow your data to be used for product improvement. The calculator models paid usage only.