Gemini 3.8 Flash Price Doubles Jan 1, 2027: What It Means for Bills
Gemini 3.8 Flash goes from $0.75/$3.75 to $1.50/$7.50 per 1M tokens on Jan 1, 2027. Worked bill examples, alternatives, and what to do before January.
- Google’s pricing page lists Gemini 3.8 Flash (paid tier, standard) at $0.75 input / $3.75 output per 1M tokens through December 31, 2026, and $1.50 / $7.50 starting January 1, 2027.
- Every other line doubles too: cached input $0.075 → $0.15, cache storage $0.50 → $1.00 per 1M tokens per hour, and the Batch, Flex and Priority rates.
- Gemini 3.7 Flash and Gemini 3.6 Flash carry the identical schedule, so moving between those three does not avoid the increase.
- Worked example: a workload of 400M input and 80M output tokens a month goes from $600 to $1,200. Batch processing brings it back to $600; caching can cut the input share further.
- Before January: measure real token usage per feature, turn on caching and batch where they fit, and test one or two cheaper alternatives on your own prompts.
Google’s Gemini API pricing page (last updated 2026-10-01 when we read it) shows a dated price change for Gemini 3.8 Flash: every per-token rate doubles on January 1, 2027. This post sets out exactly what the page lists, what a doubling does to a monthly bill, how it compares with other models in the same price range, and the practical steps worth taking in the three months before it lands. We only report what Google’s page states. The page gives no reason for the change, and we don’t guess at one.
What Google’s pricing page lists
All prices are USD per 1M tokens on the paid tier. The free tier lists Gemini 3.8 Flash input and output as free of charge and is not affected by the table below.
| Gemini 3.8 Flash, paid tier | Through Dec 31, 2026 | From Jan 1, 2027 |
|---|---|---|
| Standard input | $0.75 | $1.50 |
| Standard output (including thinking tokens) | $3.75 | $7.50 |
| Context caching (cached input) | $0.075 | $0.15 |
| Cache storage, per 1M tokens per hour | $0.50 | $1.00 |
| Batch input / output | $0.375 / $1.875 | $0.75 / $3.75 |
| Flex input / output | $0.375 / $1.875 | $0.75 / $3.75 |
| Priority input / output | $1.35 / $6.75 | $2.70 / $13.50 |
Three details on the same page are easy to miss:
- Gemini 3.7 Flash and Gemini 3.6 Flash list the same prices and the same January 1, 2027 change. Switching among the three Flash generations doesn’t change what you pay.
- Output pricing includes thinking tokens. If you run the model with heavy reasoning, those tokens are billed at the output rate, which is also doubling.
- Gemini 3.5 Flash shows no dated change as of the page’s 2026-10-01 update. It lists $1.50 input and $9.00 output today. After January, 3.8 Flash ($1.50 / $7.50) is still cheaper than 3.5 Flash on output and equal on input. That is no guarantee 3.5 Flash’s price won’t change. It just isn’t listed.
Grounding with Google Search is listed separately (5,000 free search requests a month shared across Gemini 3.x models, then $14 per 1,000) with no dated change shown.
What doubling does to a monthly bill
These are arithmetic examples using the published prices, not measured bills. Each one uses a different monthly token volume.
| Workload (monthly tokens) | Through Dec 2026 | From Jan 2027 | Change |
|---|---|---|---|
| Small app: 50M input, 10M output | $75 | $150 | +$75 |
| Mid-size: 400M input, 80M output | $600 | $1,200 | +$600 |
| Heavy RAG: 2B input, 200M output | $2,250 | $4,500 | +$2,250 |
The ratio is exactly 2x because every rate doubles, so your own bill doubles too unless you change something. The levers below change the shape of the bill, not just its size.
Heavy RAG with 70% of input served from cache (an assumed hit rate, standard tier):
- Through Dec 2026: 1.4B cached × $0.075/M = $105, plus 0.6B uncached × $0.75/M = $450, plus output 200M × $3.75/M = $750. Total $1,305.
- From Jan 2027: $210 + $900 + $1,500 = $2,610.
Caching cuts the input share sharply, but output is now more than half the bill, and output can’t be cached.
Heavy RAG moved entirely to batch (no caching): from January 2027, 2B × $0.75/M + 200M × $3.75/M = $2,250. That is exactly today’s standard-rate bill. If part of your traffic can wait for a batch job, that part costs in 2027 what it costs now.
To model your own numbers, enter them in the LLM API cost calculator, once at each price.
How it compares with same-tier alternatives
Our verified price list puts Gemini 3.8 Flash in the “mid” tier. This table uses the mid-size example (400M input, 80M output a month) at each model’s listed standard rate, plus a few small-tier models that are often tried as replacements.
| Model (tier in our list) | Input / output per 1M | Mid-size example |
|---|---|---|
| Gemini 3.8 Flash, through Dec 2026 (mid) | $0.75 / $3.75 | $600 |
| Gemini 3.8 Flash, from Jan 2027 (mid) | $1.50 / $7.50 | $1,200 |
| Gemini 3.5 Flash (mid) | $1.50 / $9.00 | $1,320 |
| Gemini 2.5 Flash (mid) | $0.30 / $2.50 | $320 |
| Claude Sonnet 5.5 (mid) | $2.00 / $10.00 | $1,600 |
| GPT-5.6 Terra (mid) | $2.00 / $12.00 | $1,760 |
| GPT-5.4 mini (small) | $0.75 / $4.50 | $660 |
| Claude Haiku 4.5 (small) | $1.00 / $5.00 | $800 |
| Gemini 3.5 Flash-Lite (small) | $0.30 / $2.50 | $320 |
| Gemini 3.1 Flash-Lite (small) | $0.25 / $1.50 | $220 |
| GPT-5.6 Luna (small) | $0.20 / $1.20 | $176 |
Read this table with three caveats:
- Token counts are not portable. The table assumes the same 400M / 80M tokens on every model. A different tokenizer counts the same text differently. Anthropic says its Claude 4.7 and later models produce about 30% more tokens than earlier ones for the same text. Measure your prompts on each candidate before trusting a comparison (why counts differ).
- Price is not capability. A small-tier model that costs a fifth as much is only cheaper if it does the job. A model that needs more retries, longer prompts or more output to reach the same quality can end up costing more.
- Older models. Gemini 2.5 Flash is an older generation that is still listed on the pricing page at $0.30 / $2.50. Confirm it’s available to your project, and check its deprecation status, before planning around it.
Even after January, Gemini 3.8 Flash stays below the mid-tier Claude and GPT models in this list and slightly below Gemini 3.5 Flash. The doubling matters most to teams who picked it as a cost-efficient default and now find small-tier models closer in price.
What to do before January 1
1. Measure where the tokens go. Break usage down by feature or endpoint: input vs output, cached vs uncached, thinking tokens. The response usage metadata reports these, and Google’s token guide covers how to count tokens before sending. A doubling of a $75 line item is a footnote. A doubling of a $4,500 one is a project.
2. Turn on caching deliberately. Implicit caching is on by default for Gemini 2.5 and newer, but Google describes it as having no cost-saving guarantee. The minimum for Gemini 3.8 Flash is 4,096 tokens. For a large shared context, an explicit cache gives guaranteed cached-token pricing, but adds storage: a 1M-token cache kept for a 720-hour month is $360 now and $720 from January. The prompt caching guide walks through the structure that makes hits likely.
3. Move non-urgent work to batch. Batch is 50% off and has a 24-hour target turnaround (jobs expire after 48 hours). Back-fills, evaluations and nightly processing are the obvious candidates.
4. Trim output and thinking. Output is the larger rate and includes thinking tokens. Tighter output formats (JSON fields instead of prose) and lower reasoning settings where quality allows reduce the part of the bill that caching can’t touch.
5. Run a bake-off on your own data. Pick 100–500 real requests, run them on Gemini 3.8 Flash and two candidates from the table above, and compare quality, measured token counts and total cost at January 2027 prices. Decide per feature. A classifier might move to a small-tier model while a reasoning-heavy feature stays.
6. Put a ceiling on spend. A doubling is predictable. A runaway loop is not. A per-run cost cap limits the damage either way.
7. Re-check the pricing page in December. We report the schedule as Google lists it on 2026-10-01. Vendors do revise announced prices, so confirm the rates before you commit budget or switch models.