Choosing an LLM API Model by Cost Tier: Small, Mid, Flagship, Frontier

A practical framework for picking an LLM API tier by cost: worked examples for a classifier, RAG chatbot and long-form writer, plus routing and escalation.

By the ailog editors · Published Oct 1, 2026 · 7 min read · How we work
In short
  • In our verified price list, input prices run from $0.10 per 1M tokens (small models) to $10 (frontier), a 100x spread. Output runs from $0.40 to $50.
  • Which price matters depends on the workload’s input/output ratio. A classifier is about 93% input cost. A long-form writer is roughly 90% output cost.
  • Tier labels don’t track price exactly. GPT-6.1 Sol (flagship) and Claude Sonnet 5.5 (mid) both list $2 input / $10 output.
  • Start with the cheapest tier that passes your own quality test, then add routing: a small model first, with escalation to a larger one only when needed.
  • A more expensive model can be cheaper overall if it finishes in fewer steps, retries less or writes less. Measure cost per completed task, not per token.

Picking a model by its per-token price is like picking a car by the price of fuel. It matters, but how far you drive matters more. This guide gives a practical framework for choosing an LLM API tier by cost: what the tiers look like in today’s published prices, how the input/output ratio of your workload changes which number matters, three worked examples, and how routing lets you use expensive models only where they pay for themselves. Every price below comes from the providers’ official pricing pages as recorded in our LLM API pricing table on 2026-10-01.

The four tiers in today’s prices

Our price list groups models into four tiers. The ranges, in USD per 1M tokens on standard short-context rates:

Tier Examples Input Output
Small GPT-6 Luna, GPT-5.6 Luna, Gemini 3.1 Flash-Lite, Gemini 2.5 Flash-Lite, Claude Haiku 4.5 $0.10 – $1.00 $0.40 – $5.00
Mid Gemini 3.8 Flash, Gemini 3.5 Flash, Claude Sonnet 5.5, GPT-5.6 Terra $0.30 – $2.00 $2.50 – $12.00
Flagship GPT-6.1 Sol, Gemini 3.1 Pro Preview, Claude Opus 5.5, GPT-5.6 Sol $1.25 – $4.00 $10.00 – $20.00
Frontier Claude Fable 5.1, GPT-6 Astra $10.00 $50.00

Things to know before using this table:

  • Tiers overlap. Claude Haiku 4.5 (small, $1 / $5) costs more than Gemini 3.8 Flash (mid, $0.75 / $3.75). GPT-6.1 Sol (flagship) and Claude Sonnet 5.5 (mid) are both $2 / $10. Compare prices, not labels.
  • Output costs more than input everywhere, from 4x to a little over 8x the input price across this list.
  • Some prices are temporary. GPT-5.6 Sol’s price is promotional through at least November 21, 2026, according to OpenAI. Gemini 3.8 Flash doubles on January 1, 2027 (details).
  • Long prompts cost more on some models. OpenAI’s short-context rates apply up to 272K input tokens, and Gemini Pro-class rates here are for prompts up to 200k tokens. All examples below stay well under both.
  • Tokens aren’t equal across models. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text than earlier Claude models. The examples below assume equal token counts for readability. Your real comparison should use measured token counts.

Input/output ratio decides which price matters

Every request has an input side (instructions, context, conversation) and an output side (the answer; Google’s pricing page notes that Gemini output prices include thinking tokens). Because output costs 4–8x more per token, the shape of the workload determines which column dominates. The three examples below are arithmetic examples using published prices, with assumed volumes.

1. Ticket classifier. 1,000,000 requests a month, each 400 input tokens and 5 output tokens (a label). That’s 400M input and 5M output, and input is about 93–94% of the cost on every model.

Model (tier) Monthly cost
GPT-6 Luna (small) $42.50
GPT-5.6 Luna (small) $86.00
Gemini 3.1 Flash-Lite (small) $107.50
Gemini 3.8 Flash (mid) $318.75
Claude Haiku 4.5 (small) $425.00
Claude Sonnet 5.5 (mid) $850.00
Claude Opus 5.5 (flagship) $1,700.00
Claude Fable 5.1 (frontier) $4,250.00

For a classifier, the input price is everything, and the gap between the cheapest small model and a frontier model is 100x. Classification is also where small models are most often good enough, and where prompt caching (a long fixed instruction block) and batch processing help most.

2. RAG support chatbot. 200,000 requests a month, each 6,000 input tokens (instructions plus retrieved passages) and 400 output tokens. That’s 1.2B input and 80M output. Input is still 71–75% of the cost.

Model (tier) Monthly cost
GPT-6 Luna (small) $160
GPT-5.6 Luna (small) $336
Gemini 3.8 Flash (mid) $1,200
Claude Haiku 4.5 (small) $1,600
GPT-6.1 Sol (flagship) / Claude Sonnet 5.5 (mid) $3,200
Claude Opus 5.5 (flagship) $6,400
GPT-6 Astra (frontier) $16,000

Here caching is the big lever, because the instruction block repeats on every request. A 10x cheaper cached-input rate on that part of the prompt changes the comparison more than moving one tier down. See the prompt caching guide for the mechanics.

3. Long-form writer. 5,000 articles a month, each 3,000 input tokens (brief and outline) and 4,000 output tokens. That’s 15M input and 20M output, and output is now 87–89% of the cost.

Model (tier) Monthly cost
GPT-6 Luna (small) $11.50
Gemini 3.8 Flash (mid) $86.25
Claude Haiku 4.5 (small) $115.00
Claude Sonnet 5.5 (mid) $230.00
Gemini 3.1 Pro Preview (flagship) $270.00
Claude Opus 5.5 (flagship) $460.00
Claude Fable 5.1 / GPT-6 Astra (frontier) $1,150.00

Output-heavy work is cheap in absolute terms at moderate volumes, so this is where a higher tier is easiest to justify if quality visibly improves. To save money here, cut output: tighter length limits, structured outlines, fewer drafts. Caching does little.

Run your own volumes through the LLM API cost calculator. It applies the same published prices.

Routing and escalation

You don’t have to pick one model for everything. Routing sends each request to the cheapest model likely to handle it, and escalation retries on a larger model when the first answer fails a check.

An arithmetic example using the RAG chatbot above, with an assumed escalation rate:

  • All requests on GPT-5.6 Luna: $0.00168 per request, $336 a month.
  • All requests on GPT-6.1 Sol: $0.016 per request, $3,200 a month.
  • Every request goes to GPT-5.6 Luna first, and 20% are escalated to GPT-6.1 Sol (paying for both calls): $0.00168 + 0.2 × $0.016 = $0.00488 per request, $976 a month.

Routing costs about 30% of the all-flagship bill and puts the flagship model on the hardest fifth of traffic. The escalation rate is the number that decides everything. At 50% escalation, the routed cost per request is $0.00968, still well under the flagship-only $0.016. Escalation pays as long as the small-model call costs less than the flagship calls it avoids.

Practical ways to decide when to escalate:

  • Rule-based: route by request type, length (count tokens first with the provider’s counting endpoint), or customer plan.
  • Validation-based: escalate when the cheap model’s output fails a schema check, a confidence threshold, a test suite, or a grader.
  • User-based: offer “try harder” as an explicit action rather than spending flagship tokens by default.

Log which tier served each request and whether it was escalated. Without that, you can’t tell whether routing is saving money or quietly sending everything to the expensive model.

When the expensive model is cheaper overall

Per-token price is only one factor. Cost per completed task also depends on how many tokens the model uses to get there. This hypothetical example uses assumed step counts, not observed behaviour, to show the mechanism:

  • A coding agent on Claude Sonnet 5.5 takes 30 turns at about 40,000 input tokens each, plus 30,000 output tokens in total: 1.2M × $2/M + 30K × $10/M = $2.70 per task.
  • Suppose Claude Opus 5.5 finishes the same task in 12 turns, with 15,000 output tokens: 480K × $4/M + 15K × $20/M = $2.22 per task.

If the larger model really needs fewer than half the turns, it’s cheaper despite double the per-token price. If it needs 20 turns, it isn’t. Only measurement on your own tasks tells you which case you’re in. The same logic applies to:

  • Retries and failures. A small model that fails a third of the time and needs a retry, or a human fix, can cost more per success.
  • Prompt overhead. A weaker model may need long few-shot examples that a stronger model doesn’t. Those are paid on every call.
  • Verbosity. A model that answers in 150 tokens instead of 400 saves on the expensive side of the bill.

Whichever tier you run, a per-run cost cap stops one long agent loop from erasing a month of savings.

A selection checklist

  1. Write down the workload shape: requests per month, input and output tokens per request, and latency needs.
  2. Shortlist two or three models across tiers from the pricing table, including at least one small model.
  3. Build a 100–300 item test set from real traffic and define pass/fail before you look at outputs.
  4. Measure tokens and quality on each candidate, using each provider’s token counting endpoint for input and real runs for output.
  5. Compute cost per successful task, including retries and escalations, in the cost calculator.
  6. Apply discounts that fit: caching for repeated prefixes, and the Batch API’s 50% discount for anything that can wait.
  7. Re-run the comparison when prices change. Promotional prices end and scheduled increases arrive. GPT-5.6 Sol’s price is promotional and Gemini 3.8 Flash has a dated increase, so both are worth re-checking.
Sources
  1. Anthropic: Pricing
  2. OpenAI: Pricing
  3. Google: Gemini Developer API pricing
  4. Anthropic: Token counting
  5. OpenAI: Counting tokens

Related