Why LLM Token Counts Differ Between Models (and How to Measure Them)
Same text, different token counts: tokenizer changes like Claude 4.7+'s ~30% more tokens, non-English text, code, and official token counting endpoints.
- A token is a unit of the model’s vocabulary, not a fixed number of characters. Each provider, and sometimes each model generation, splits the same text differently.
- Anthropic says Claude 4.7 and later models use a newer tokenizer that produces roughly 30% more tokens for the same text than Claude Sonnet 4.6 and earlier. Its advice: recount prompts on the model you plan to use.
- All three providers offer an API to count input tokens before sending: Anthropic
/v1/messages/count_tokens(documented as free, with its own rate limits), OpenAI/v1/responses/input_tokens, Googlemodels/…:countTokens. - In our own countTokens check on four Gemini models, a Korean sentence used 22 tokens against 13 for its English counterpart, and a short JSON object used 37 tokens for 72 characters.
- Cost comparisons based on price per million tokens are only valid if you multiply by tokens measured on each model. Price × measured tokens is the real cost.
Every LLM price list is quoted per million tokens, so it’s tempting to compare models by reading the price column. That works only if a million tokens means the same amount of text everywhere, and it doesn’t. The same prompt can produce noticeably different token counts on two providers, or on two generations of one provider’s models. This explainer covers why, what the providers themselves say about it, how to count tokens with their official endpoints, and what changes in a cost comparison once you measure.
What a token is, and why counts differ
A tokenizer splits text into pieces from a fixed vocabulary. Google’s token guide explains that a token can be a single character or a whole word, and that long words are broken into several tokens. Which pieces exist in the vocabulary is a design choice of each model family. A common English word may be one token in one vocabulary and two in another. A Korean syllable, a run of spaces in code, or a JSON quote-and-colon sequence may be split very differently.
Providers give rough rules of thumb, and they are only that:
- Google: for Gemini models a token is about 4 characters, and 100 tokens is about 60–80 English words.
- Anthropic (in its web fetch pricing notes): an average 10 kB web page is about 2,500 tokens, and a 100 kB documentation page about 25,000.
Rules of thumb like these are calibrated on typical English. They drift quickly for other languages, code and structured data, as the measurements below show.
Tokenizers change between model generations
The clearest official statement comes from Anthropic’s pricing page. Claude 4.7 and later models use a newer tokenizer, and in Anthropic’s words:
This tokenizer produces approximately 30% more tokens for the same text.
The page adds that the exact increase depends on the content and the shape of the workload, and that Claude Sonnet 4.6 and earlier use the previous tokenizer. Anthropic’s token counting docs draw the practical conclusion: billing on these models reflects the new tokenizer’s counts, so don’t reuse token counts measured on an older model to estimate costs or context-window fit. Recount with the model ID you plan to use.
This has a direct price consequence. Here is an arithmetic example using Anthropic’s approximate figure, not a measurement:
- Claude Sonnet 4.6 lists $3 per 1M input tokens. Claude Sonnet 5.5 lists $2.
- Suppose a prompt measures 1,000,000 tokens on Sonnet 4.6. If it comes out about 30% larger on Sonnet 5.5, that is ~1,300,000 tokens × $2 = ~$2.60, against $3.00 on Sonnet 4.6.
- The newer model is still cheaper for this input, but by about 13%, not the 33% the price column suggests. The real gap depends on your content, which is why Anthropic tells you to recount.
Hidden overhead changes between generations too. Anthropic’s pricing page lists the tokens added by the tool-use system prompt per model. With tool_choice auto or none, it’s 286 tokens on Claude Opus 5.5, 290 on Opus 4.8, 675 on Opus 4.7 and 497 on Opus 4.6. That’s small per request, but it’s billed on every request that includes tools.
Counting tokens with the official endpoints
Local tokenizer libraries are fine for rough English estimates. They can’t see what the API adds. OpenAI’s guide lists the gaps: local tokenizers like tiktoken don’t handle images and files, and tools and schemas add tokens that are hard to count locally. Its counting endpoint also includes the formatting tokens that represent request structure, such as message roles and boundaries. All three providers expose a counting call that takes the same payload as a real request.
Anthropic (claude-opus-5-5 is the model ID used in Anthropic’s own docs):
curl https://api.anthropic.com/v1/messages/count_tokens \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model": "claude-opus-5-5",
"system": "You are a scientist",
"messages": [{"role": "user", "content": "Hello, Claude"}]}'
Anthropic documents this endpoint as free to use, with requests-per-minute limits by usage tier that are separate from message creation. It returns an estimate that can differ slightly from actual usage, and it may include system-added tokens that you are not billed for. It doesn’t accept server tools (web search, code execution and others) or URL-sourced images and documents. Send media as base64 to count them.
OpenAI:
curl https://api.openai.com/v1/responses/input_tokens \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-6-astra", "input": "Tell me a joke."}'
The input token count endpoint takes the same input format as the Responses API, including messages, images, files, tools and conversations. OpenAI’s guide doesn’t state a price for it, so check your usage dashboard if that matters to you.
Google Gemini:
curl -X POST "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:countTokens" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"contents": [{"parts": [{"text": "The quick brown fox."}]}]}'
Google’s guide also documents fixed rates for media: images up to 384 pixels in both dimensions count as 258 tokens, larger images are tiled into 768×768 tiles at 258 tokens each, video is 263 tokens per second (static processing), and audio 32 tokens per second.
After a request, use the usage fields in the response as the source of truth for billing. Counting endpoints are for planning, routing and staying under context limits.
Our measurement: languages, code and JSON on Gemini
We ran the Gemini countTokens endpoint (v1beta/models/{model}:countTokens) on 2026-10-01 with five short strings, on four models: gemini-3.8-flash, gemini-3.5-flash, gemini-3.1-flash-lite and gemini-3.1-pro-preview. The exact inputs:
en: The invoice total is due within thirty days of the delivery date.
ko: 청구서 총액은 배송일로부터 30일 이내에 지불해야 합니다.
ja: 請求書の合計金額は、納品日から30日以内にお支払いください。
code: def total(items):\n return sum(i.price * i.qty for i in items)\n
json: {"invoice_id": 1042, "total": 129.50, "currency": "USD", "due_days": 30}
(In the code string, \n stands for a real newline and the indentation is four spaces.)
| String | Characters | Tokens (each of the four models) |
|---|---|---|
| English sentence | 65 | 13 |
| Korean sentence (a translation of the English) | 32 | 22 |
| Japanese sentence (a translation of the English) | 30 | 17 |
| Python function | 65 | 23 |
| JSON object | 72 | 37 |
What this small check does and doesn’t show:
- All four Gemini models returned identical counts for these five strings. That is consistent with a shared tokenizer for these strings, but five strings don’t prove the tokenizers are identical.
- The Korean and Japanese sentences are translations, not the same text, so they don’t give a general “language multiplier”. They do show that “4 characters per token” can be badly off. The Korean sentence is about 1.5 characters per token, and the English one 5.
- The JSON object averaged under 2 characters per token. Numbers, quotes and punctuation are expensive relative to their length. If you send large JSON payloads, measure them.
- We didn’t run the same strings through Anthropic’s or OpenAI’s endpoints for this article, so we make no cross-provider claim from these numbers.
Why cost comparisons need measured tokens
The cost of a task is price per token × tokens the task uses, summed over input, cached input and output. Price lists give you the first factor. Only measurement gives you the second, and the second varies by:
- Tokenizer: across providers and across generations (Anthropic’s ~30% statement).
- Language and content type: prose vs code vs JSON vs non-English text.
- Hidden overhead: tool-use system prompts, message formatting, tool schemas.
- Output behaviour: one model may answer in 150 tokens where another uses 400. Thinking tokens add to this; Google’s pricing page, for example, states that Gemini output prices include thinking tokens.
The practical method:
- Take 50–200 real requests from your logs, including non-English and structured ones if you have them.
- Count input tokens with each candidate model’s counting endpoint.
- Run a sample for real to measure output tokens, which no counting endpoint can predict.
- Multiply by each model’s prices from the LLM API pricing table, or enter the totals in the cost calculator.
That’s the same method we recommend when choosing a model tier by cost or weighing alternatives to Gemini 3.8 Flash after its January 2027 price change.
Checklist
- Recount token budgets whenever you change model generation, not just provider.
- Use the provider’s counting endpoint for planning and the response
usagefields for billing reconciliation. - Treat “4 characters per token” as English-only, and only roughly.
- Budget separately for tool definitions and structured output schemas. They are sent, and billed, on every call.
- Watch prompt caching minimums (for example 4,096 tokens on Gemini 3.8 Flash and Claude Haiku 4.5). They are in tokens too, so a tokenizer change can move a prompt above or below the threshold. See how prompt caching works.
- Cap spend per run so an unexpectedly token-heavy input can’t surprise you: per-run cost caps.