Prompt Caching Explained: How Claude, GPT and Gemini Bill Cache Hits

How prompt and context caching works at Anthropic, OpenAI and Google: minimum sizes, TTLs, write premiums, storage fees, and worked cost examples.

By the ailog editors · Published Oct 1, 2026 · 9 min read · How we work
In short
  • All three providers charge much less for input tokens that come out of a cache: 10% of the base input price for most current models, and 5% or less on a few (Claude Opus 5.5, Claude Fable 5.1, GPT-6.1 Sol).
  • Anthropic and OpenAI (GPT-5.6 and later) bill a cache write at 1.25x the base input price. Google’s implicit cache has no write fee. Google’s explicit cache instead charges storage per token per hour.
  • Minimum cacheable prefixes differ: 512 to 4,096 tokens at Anthropic depending on the model, 1,024 at OpenAI for GPT-5.6 and later, 2,048 to 4,096 at Google.
  • Default lifetimes differ too: 5 minutes at Anthropic (1 hour optional), at least 30 minutes on newer OpenAI models, and a TTL you choose (1 hour by default) for Gemini explicit caches.
  • Caching only matches an identical prefix. Put stable content first and anything that changes per request last, or you pay for writes and never get reads.

Prompt caching can cut the input side of an LLM bill by more than half when requests share a long prefix. The idea is the same everywhere. If many requests start with the same long block of tokens (system instructions, tool definitions, a reference document, a long conversation history), the provider can keep the processed state of that block and bill later reads at a fraction of the normal input price. The billing rules differ, and a few of the differences decide whether caching saves 70% or costs you extra. Everything below comes from each provider’s own documentation, read on 2026-10-01. Prices come from our verified price list.

What a cache hit actually is

A cache entry is keyed on a prefix: the exact tokens from the start of the request up to some boundary (Anthropic and OpenAI call it a breakpoint). A later request gets a hit only if its own tokens are identical up to that boundary. One changed character early in the prompt (a timestamp in the system prompt, a reordered tool list, a different user ID) changes everything after it. That part of the request then becomes a miss.

The order in which the request is assembled matters:

  • Anthropic builds the prefix in the order tools, then system, then messages. Changing tool definitions invalidates the whole cache. Changing the system prompt invalidates the system and message caches.
  • OpenAI caches the whole rendered context the model sees, including its own hidden instructions, developer messages, tool definitions and conversation history. Its guide lists settings that change the cached prefix, among them tools, text.format (structured outputs), reasoning.effort and text.verbosity.
  • Google says cached content is simply a prefix to the prompt, and recommends putting large, commonly reused content at the beginning.

Caches are also scoped. Anthropic isolates caches per workspace on the Claude API (per organization on Bedrock and Google Cloud). OpenAI caches are never shared across organizations or across regional processing boundaries.

The three billing models side by side

Anthropic (Claude) OpenAI (GPT-5.6 and later) Google (Gemini)
How it turns on Opt-in: top-level cache_control (automatic) or cache_control on blocks (explicit) On by default (implicit); explicit breakpoints optional Implicit on by default for Gemini 2.5 and newer; explicit caches created via API
Cache read price 0.1x base input (0.05x Opus 5.5; 0.025x Fable 5.1) 0.1x base input (0.05x GPT-6.1 Sol) Listed per model, for example $0.075 vs $0.75 on Gemini 3.8 Flash
Write premium 1.25x (5-minute) or 2x (1-hour) 1.25x None for implicit
Storage fee None None Explicit caches only, per 1M tokens per hour
Minimum prefix 512–4,096 tokens by model 1,024 visible input tokens 2,048–4,096 by model
Lifetime 5 min default, 1 hour optional, refreshed on each hit At least 30 min after last write or reuse Implicit: not stated. Explicit: TTL you set, default 1 hour

A few details from the docs matter more than the table suggests:

  • Anthropic minimums by model. 512 tokens for Claude Fable 5.1, Opus 5.5, Opus 5 and Sonnet 5.5. 1,024 for Sonnet 5, Opus 4.8 and Sonnet 4.6. 4,096 for Haiku 4.5. A prompt below the minimum is processed without caching, and no error is returned. The only signal is that cache_creation_input_tokens and cache_read_input_tokens are both 0.
  • Anthropic lifetime is measured from the start of the request. If a response takes 4 minutes to stream, the follow-up has about 1 minute left on a 5-minute cache. A cache entry only becomes available once the first response begins, so parallel requests fired at the same moment all miss.
  • OpenAI write pricing is a replacement, not an extra fee. The guide says each input token is billed at either the uncached, cached or cache-write rate. On GPT-5.6 and later, OpenAI’s own arithmetic: one write plus one full reuse costs 1.35x the plain input cost, against 2x for sending it twice uncached.
  • OpenAI models before GPT-5.6 have no write charge and use prompt_cache_retention (in_memory or 24h) instead of the newer prompt_cache_options.ttl, whose only supported value is 30m.
  • Google implicit caching comes with no cost-saving guarantee. Explicit caching is the option Google describes as having a cost-saving guarantee. It is currently in Beta and available through the generateContent API, not the Interactions API.

Worked example: a support bot with a 10,000-token prefix

This is an arithmetic example using published prices, not a measured bill. The assumptions:

  • 100,000 requests a month.
  • Each request is a 10,000-token static prefix (instructions plus reference docs), a 500-token user message that changes every time, and a 300-token answer.
  • 95% of requests find the prefix in cache and 5% have to write it again (the cache expired during a quiet period). This hit rate is an assumption; yours depends on traffic.
  • Token counts are treated as identical across vendors to keep the arithmetic readable. In reality they differ. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text (see why token counts differ between models).
Model and setup No caching With caching
Claude Sonnet 5.5, explicit breakpoint at end of prefix, 5-min TTL $2,400 $715
GPT-6.1 Sol, explicit-only mode, breakpoint at end of prefix $2,400 $620
Gemini 3.8 Flash, implicit cache, 95% hits assumed $900 $258.75
Gemini 3.8 Flash, one explicit cache kept all month $900 $228.60 + cache creation

How the Claude line breaks down: 95,000 reads × 10,000 tokens × $0.20/M = $190. 5,000 writes × 10,000 × $2.50/M = $125. The varying 500 tokens at $2/M = $100. Output 30M tokens × $10/M = $300. Total $715.

The GPT-6.1 Sol line uses explicit-only mode: one breakpoint at the end of the static prefix, so the changing 500 tokens are billed at the normal input rate with no write charge. We don’t price OpenAI’s implicit mode for this single-turn shape. Implicit mode places its breakpoint at the end of the latest user message, which here is the part that changes every request, and OpenAI’s own example for a stable-prefix, changing-suffix prompt uses explicit-only mode.

The Gemini explicit line assumes the cache lives all month. Storage is 10,000 tokens × $0.50 per 1M tokens per hour × 720 hours = $3.60. Reads are 100,000 × 10,000 × $0.075/M = $75. The pages we read do not state what creating the cache costs, so that is left out. Storage scales with size: a 1M-token cache kept for a 720-hour month is $360 in storage on Gemini 3.8 Flash today, and $720 from January 1, 2027, when Google’s listed storage price doubles (details here).

To try your own volumes, plug the cached and uncached token counts into the LLM API cost calculator.

When a cache write does not pay for itself

Anthropic’s documentation gives the break-even directly. A 5-minute write (1.25x) pays off after one cache read. A 1-hour write (2x) pays off after two. The real risk is not the break-even point. It is writing caches that are never read:

  • Low or bursty traffic. At one request every 10 minutes, a 5-minute cache expires between requests and every call pays the 1.25x premium. Either use the 1-hour TTL (2x write, so you need at least two reads per hour to win) or accept no caching.
  • Breakpoint on the changing block. Anthropic’s docs call this out as a common mistake. If the block you mark contains a timestamp or the user’s message, the hash changes every request and the system writes a new entry every time without ever reading one. Its automatic caching puts the breakpoint on the last cacheable block, so a single-turn prompt whose last block varies falls into the same trap. For that shape, use an explicit breakpoint at the end of the static part.
  • Prefixes just under the minimum. OpenAI’s guide works through this. With a 1,024-token minimum, 0.1x reads and 1.25x writes, expanding a shorter prefix up to 1,024 tokens becomes cheaper over 10 requests once the original is at least 221 tokens long. Below about 102 tokens it never pays. Anthropic makes the same point: padding a prompt up to the threshold with useful stable content is often worthwhile.
  • Gemini explicit caches that sit idle. Storage is charged per hour whether or not anyone reads the cache. A large cache with a long TTL and few reads can cost more than sending the tokens uncached.

How to structure prompts for cache hits

The rules are the same at every provider: stable first, variable last, and never rewrite history.

[tools / function definitions]      <- identical every request, never reorder
[system / developer instructions]   <- no timestamps, user names or request IDs
[reference documents, examples]     <- stable for hours or days
---------- cache breakpoint here ----------
[retrieved snippets for this query] <- changes per request
[conversation turns, appended only]
[current user message]

Provider-specific moves:

  • Anthropic: put cache_control on the last block that is identical across requests. You can use up to 4 breakpoints, for example one after tools and one after the system prompt. In long conversations, note the 20-block lookback window. If a conversation grows more than 20 blocks past the last write, add a second breakpoint earlier. To add instructions mid-conversation without breaking the cache, newer models accept a role: "system" message appended inside messages instead of an edited top-level system.
  • OpenAI: keep tool lists stable. Disable tools per request with tool_choice: "none" or allowed_tools rather than removing definitions. On GPT-6 models, change reasoning effort with an appended configuration_update item instead of the top-level setting. On models before GPT-5.6, use a stable prompt_cache_key so related requests route to the same machine.
  • Gemini: for implicit hits, put large shared content first and send similar requests close together in time. For guaranteed savings on a large shared document, create an explicit cache with a TTL that matches how long you actually use it.

Checklist: verify that caching is working

  1. Read the usage fields on every response. Anthropic: cache_creation_input_tokens, cache_read_input_tokens, input_tokens. OpenAI: usage.input_tokens_details.cached_tokens and cache_write_tokens. Gemini: the cached-token count in the usage metadata.
  2. Compute a token hit rate. Divide cached tokens by total input tokens, per day and per endpoint. OpenAI recommends exactly this metric.
  3. Grep your prompt builder for anything dynamic before the breakpoint: dates, request IDs, user names, randomly ordered lists, JSON serialized with unstable key order.
  4. Check the model’s minimum. A 3,000-token prefix caches on Claude Sonnet 5.5 but silently does nothing on Claude Haiku 4.5 (4,096 minimum) or Gemini 3.8 Flash (4,096).
  5. Combine with batch where latency allows. Anthropic says caching multipliers stack with the 50% batch discount, though cache hits inside batches are best-effort. See when the Batch API’s 50% discount is worth it.
  6. Cap the downside. Caching lowers average cost but does not bound a runaway loop. A per-run cost cap does.

Prices and rules change. Check the provider pages in the sources list before you commit to a design.

Sources
  1. Anthropic: Pricing
  2. Anthropic: Prompt caching
  3. OpenAI: Pricing
  4. OpenAI: Prompt caching
  5. Google: Gemini Developer API pricing
  6. Google: Context caching (Gemini API)
  7. Google: Context caching for generateContent (implicit and explicit)

Related