A Per-Run Cost Cap for LLM API Calls: A $0.20 Hard Limit in Code

How our blog pipeline prices every Gemini response from usage metadata, stops at $0.20 per run, and what five real runs cost (about $0.09–$0.11 each).

By the ailog editors · Published Oct 1, 2026 · 7 min read · How we work
In short
  • Every Gemini call in our blog pipeline goes through one wrapper. It prices the response from usageMetadata, adds it to a running total, and throws once the run passes $0.20.
  • Five complete runs cost $0.0863–$0.1124 each, about $0.10. Text was only $0.015–$0.041 of that. The single AI hero image (around $0.072 as we log it) was 64–83% of every run.
  • Our plan had estimated about $0.005 of text per post. Measured text cost was 3–8× higher, mostly because one post takes up to four model calls, not one.
  • The cap bounds overspend to one call, not zero. The check runs after the money is spent, and a call can start just under the line.
  • Gemini 3.8 Flash text prices double on January 1, 2027. Our most expensive run would go from $0.112 to about $0.153 (assuming image prices stay the same). That’s still under the cap, with less room.

Our blog pipeline calls a paid model several times per post. It’s built for our client-facing blog, a small Korean web studio’s blog, and its architecture is described here. The plan is to run it unattended every few hours. A loop bug, a prompt that explodes the output, or a retry gone wrong could run up a bill while nobody is watching. So before scheduling anything, we put a hard per-run limit into the one function every model call passes through.

The prices we price against

The pipeline uses two models. These are the paid-tier prices we use, checked against Google’s pricing page on October 1, 2026:

Model Input Output Image output Source note
Gemini 3.8 Flash (text) $0.75 / 1M tokens $3.75 / 1M tokens — Through Dec 31, 2026; output “including thinking tokens”
Gemini 3.8 Flash from Jan 1, 2027 $1.50 / 1M tokens $7.50 / 1M tokens — Announced price change
Gemini 3.1 Flash Image $0.50 / 1M (text/image) $3 / 1M (text and thinking) $60 / 1M tokens A 1K image is 1,120 tokens, “equivalent to $0.067 per image”

At $0.067 a picture, the image model is by far the biggest cost line. That shaped the design: only one AI image per post, with every other image rendered from HTML for free. (How the free images work.)

The wrapper: price every response, then check the total

All model traffic goes through gemini() in lib.mjs. It does three things: it refuses to start once the run is over budget, it makes the request, and it charges the response:

const PRICE = {
  'gemini-3.8-flash': { in: 0.75, out: 3.75 },
  'gemini-3.1-flash-image': { image: 0.067, in: 0.5, out: 3.0 },
};
export const RUN_CAP_USD = 0.20;
export const cost = { usd: 0, calls: [] };

function charge(model, usage, images = 0) {
  const p = PRICE[model] || { in: 1, out: 5 };
  const inTok = usage?.promptTokenCount || 0;
  const outTok = (usage?.totalTokenCount || 0) - inTok;
  const usd = (inTok * (p.in || 0) + Math.max(outTok, 0) * (p.out || 0)) / 1e6 + images * (p.image || 0);
  cost.usd += usd;
  cost.calls.push({ model, inTok, outTok, images, usd: +usd.toFixed(5) });
  if (cost.usd > RUN_CAP_USD) throw new Error(`run cost cap exceeded: $${cost.usd.toFixed(3)}`);
}

export async function gemini(model, body) {
  if (cost.usd > RUN_CAP_USD) throw new Error('run cost cap reached before call');
  const r = await fetch(`https://generativelanguage.googleapis.com/v1beta/models/${model}:generateContent`, {
    method: 'POST',
    headers: { 'content-type': 'application/json', 'x-goog-api-key': key() },
    body: JSON.stringify(body),
  });
  const j = await r.json();
  if (!r.ok) throw new Error(`${model} ${r.status}: ${JSON.stringify(j).slice(0, 400)}`);
  const parts = j.candidates?.[0]?.content?.parts || [];
  const images = parts.filter((p) => p.inlineData).length;
  charge(model, j.usageMetadata, images);
  return { j, parts };
}

Some details are worth explaining:

  • Output is total − prompt, not candidatesTokenCount. The API reference lists separate candidatesTokenCount and thoughtsTokenCount fields. Google bills Flash output “including thinking tokens”. Subtracting the prompt from the total catches both without depending on which fields a given response fills in.
  • Unknown models get a pessimistic default of $1 in and $5 out per million. If someone swaps in a new model without updating the table, the cap still bites, just a bit early.
  • The per-call log is the audit trail. run.json writes every call with model, token counts and dollars. That’s where all the numbers below come from.
  • The key comes from the OS keychain (security find-generic-password -s <keychain-service> -w), not from a file in the repo.

What five real runs cost

These are the five complete runs of our first test brief, a Korean payment-gateway integration guide. Each figure is the logged total of its calls:

Run Model calls Text cost Hero image (logged) Total Image share Time
1 (first version, no research step) 1 text + 1 image $0.0245 $0.0718 $0.0962 75% 45 s
2 (lint failed, one rewrite) 2 text + 1 image $0.0229 $0.0717 $0.0947 76% 61 s
3 1 text + 1 image $0.0146 $0.0718 $0.0863 83% 46 s
4 (research added, one rewrite) 4 text + 1 image $0.0407 $0.0717 $0.1124 64% 82 s
5 (research, passed first time) 3 text + 1 image $0.0277 $0.0718 $0.0995 72% 80 s

So a post costs about ten cents, and three quarters of that is one picture. The biggest single text call was the first version’s: 4,740 prompt tokens and 5,572 output tokens for $0.0245. Later the research step came in. It adds a cheap keyword-pick call (about $0.003) and a research call ($0.006–$0.008). In exchange, the writing prompt carries a compact brief instead of guesses.

Our planning document had estimated about $0.005 of text per post. The measured range was $0.015–$0.041. The gap comes from output tokens and call count. A single writing call produced between 1,866 and 5,572 output tokens, a lint retry repeats the most expensive call, and thinking tokens are billed as output. If you’re estimating a similar pipeline, multiply your single-call guess by the number of calls you’ll really make. The LLM API cost calculator handles the per-call arithmetic.

Where the cap is soft

A cap written like this is a circuit breaker, not a prepaid card. We should be precise about its limits.

It trips after the spend. charge() runs once the response has arrived. Google has already billed that call. Throwing just stops the next one, and the paid result of the call that crossed the line is thrown away with the exception.

A call can start just under the line. The pre-call check is cost.usd > RUN_CAP_USD. At $0.19 a new call is allowed. If that call is the image (about $0.07), the run ends near $0.26. The real worst case is the cap plus one maximum-size call. A stricter version would check cost.usd + worstCaseFor(model) > cap before calling. For images that’s easy, because the price per image is fixed.

The image estimate runs a little high. Our table charges every output token of an image response at the $3/M text rate, then adds the flat $0.067. Image responses reported about 1,560–1,600 output tokens. Google prices a 1K image as 1,120 tokens at $60/M, which works out to the $0.067. So the image’s own tokens are counted twice: once in the flat price and again at the text rate. The logged $0.0717 is therefore a few tenths of a cent above what the official rates give. For a cap, erring high is the safe direction, but don’t treat the log as an invoice.

It’s per process. cost lives in memory and resets each run. There’s no monthly limit in the code today. The planning document describes one, which would drop the AI image and publish text only once a monthly budget is reached. That’s still on the to-do list.

From a per-run cap to a monthly budget

Without a monthly counter in code yet, the budget is arithmetic. Here it is with our most expensive observed run ($0.1124) as the unit cost and the cap ($0.20) as the absolute worst case:

Drafts per day Typical month (at $0.1124) Worst case (every run hits $0.20)
1 about $3.37 $6.00
2 about $6.74 $12.00
6 (one every 4 hours) about $20.23 $36.00

The plan calls for generating a draft every four hours and publishing one or two a day, so the six-a-day row is the one to budget for. These figures are in US dollars at the list prices above. Your bill depends on your billing account, tax and any credits.

The January 2027 price change matters here. Text input and output prices both double. With the image price unchanged, run 4 would cost about $0.0814 in text plus $0.0717 for the image, about $0.153. Run 5 would be about $0.127. Both stay under $0.20, but the headroom on a retry-heavy run drops from about 44% to about 24%. If the image price also changes, recompute. The cap number itself might need revisiting then too, which is a good reason it lives in one exported constant.

Checklist for your own pipeline

  • Route every model call through one function. A cap that only some calls honour isn’t a cap.
  • Price from the response’s usageMetadata, and treat thinking tokens as output if your provider bills them that way.
  • Default unknown models to a pessimistic price.
  • Log each call (model, tokens, dollars) to a file next to the output, so you have real numbers instead of estimates.
  • Check spent + worst case of next call before calling, not just spent afterwards.
  • Add a persistent monthly counter before you schedule anything. A per-run cap alone doesn’t stop many runs per day.
  • Put dated price changes in your calendar. Ours doubles text prices on January 1, 2027.

For choosing the model mix in the first place, our structured-output notes cover the calls these costs come from.

Sources
  1. Gemini Developer API pricing
  2. Gemini API reference: models.generateContent (UsageMetadata)

Related