Data · updated 2026-10-01

LLM API pricing comparison (October 2026)

The current per-token prices for Anthropic, OpenAI and Google's API models in one place. Our build script re-reads each provider's pricing page and refuses to publish if a number here no longer matches.

ModelProviderInputCached inputOutputBlend*
Gemini 2.5 Flash-LiteGoogle$0.10$0.010$0.40$0.18
GPT-6 LunaOpenAI$0.10$0.010$0.50$0.20
GPT-5.6 LunaOpenAI$0.20$0.020$1.20$0.45
GPT-5.4 nanoOpenAI$0.20$0.020$1.25$0.46
Gemini 3.1 Flash-LiteGoogle$0.25$0.025$1.50$0.56
Gemini 3.5 Flash-LiteGoogle$0.30$0.030$2.50$0.85
Gemini 2.5 FlashGoogle$0.30$0.030$2.50$0.85
Gemini 3.8 FlashGoogle$0.75$0.075$3.75$1.50
GPT-5.4 miniOpenAI$0.75$0.075$4.50$1.69
Claude Haiku 4.5Anthropic$1$0.10$5$2
Gemini 3.5 FlashGoogle$1.50$0.15$9$3.38
Gemini 2.5 ProGoogle$1.25$0.13$10$3.44
Claude Sonnet 5.5Anthropic$2$0.20$10$4
GPT-6.1 SolOpenAI$2$0.10$10$4
GPT-5.6 TerraOpenAI$2$0.20$12$4.50
Gemini 3.1 Pro PreviewGoogle$2$0.20$12$4.50
Claude Opus 5.5Anthropic$4$0.20$20$8
GPT-5.6 SolOpenAI$4$0.40$20$8
Claude Fable 5.1Anthropic$10$0.25$50$20
GPT-6 AstraOpenAI$10$1$50$20

USD per 1M tokens, standard paid tier, checked against each provider's official pricing page on 2026-10-01. *Blend = 3 parts input to 1 part output, a common chat/RAG ratio — your mix will differ.

How to read these numbers

All three providers bill per million tokens, separately for what you send (input) and what the model writes back (output). Output costs roughly 4–8× as much as input than input for almost every model, so a chatbot that writes long answers costs far more than a classifier that reads long documents and returns one word.

Cached input is what you pay when the beginning of your prompt is identical to a recent request — a long system prompt, a tool list, or a document you ask several questions about. On the current models it is 90% or more cheaper than regular input. Anthropic and OpenAI also charge a one-time cache write premium (about 1.25× input on Anthropic's 5-minute cache); Google bills cached tokens plus an hourly storage fee instead.

Batch pricing is 50% off both input and output at all three providers, in exchange for results that can take up to 24 hours. Anything that does not need an answer while a user waits — nightly summaries, bulk classification, embeddings back-fills — should go through the batch endpoint.

Provider notes that change the math

Anthropic (Claude). Claude 4.7+ models use a newer tokenizer that produces roughly 30% more tokens for the same text (per Anthropic's pricing page). Cache write shown is the 5-minute write rate. Official pricing page.

OpenAI (GPT). Short-context rates. Long-context prompts are billed at higher rates (see source). GPT-5.6 Sol's price is promotional through at least Nov 21, 2026. Official pricing page.

Google (Gemini). Paid-tier text rates. Gemini 3.8 Flash rises to $1.50 input / $7.50 output on Jan 1, 2027. Pro-class prices shown are for prompts up to 200k tokens. Official pricing page.

Which tier should you start with?

Start with the cheapest "small" model that passes your own test set, then move up only the requests that fail. In our own pipelines a mid-tier Flash-class model handles drafting and structured extraction, and we reserve flagship models for steps where a wrong answer is expensive. The cost calculator lets you plug in your traffic and see the difference per month.

Methodology

Prices are the standard paid-tier rates for text, for prompts under each provider's long-context threshold. We exclude free tiers, regional/data-residency uplifts, fast/priority modes and enterprise discounts. Preview models can change price or disappear without notice. If you spot a mismatch, email [email protected] and we will correct it.