Batch API 50% Discount: When to Use It at Anthropic, OpenAI and Google
How the Claude, OpenAI and Gemini batch APIs work: the 50% discount, 24-hour windows, size limits, request formats, good-fit jobs and a monthly example.
- Anthropic, OpenAI and Google all price batch requests at 50% of their standard rates, on both input and output tokens.
- The trade is latency. All three target completion within 24 hours. Anthropic says most batches finish in under an hour; Google says most finish well within 24 hours.
- Limits differ: Anthropic allows 100,000 requests or 256 MB per batch, OpenAI 50,000 requests or 200 MB per file, and Google 2 GB per JSONL input file (20 MB for inline requests).
- Good fits: classification, extraction, summarization, embeddings, evals, and anything that runs overnight. Bad fits: anything a user is waiting on.
- Results come back unordered. Always match them by your own ID (
custom_idat Anthropic and OpenAI,keyin Gemini JSONL files).
If a job doesn’t need an answer in seconds, sending it through a batch API is the simplest way to halve its token bill. You don’t need a new model, a prompt rewrite or a quality trade-off. You need to tolerate a delay and write a little plumbing to submit, poll and collect. This guide covers how each provider’s batch API works, according to its own documentation as read on 2026-10-01, and works through what the discount means for a realistic monthly job.
What the 50% discount covers
All three providers state the discount plainly:
- Anthropic: all batch usage is charged at half the standard API price. Its pricing page lists batch rates per model. Claude Sonnet 5.5 is $1 input / $5 output per million tokens in batch, against $2 / $10 standard.
- OpenAI: batch requests cost 50% less than the synchronous APIs. The pricing page has a separate Batch table. GPT-6.1 Sol is $1.00 / $5.00 in batch, against $2.00 / $10.00 standard.
- Google: batch usage is priced at half the standard interactive rate for the same model. Gemini 3.8 Flash is $0.375 / $1.875 in batch through December 31, 2026, against $0.75 / $3.75 standard.
Discounts stack with caching in some cases. Anthropic says prompt caching and batch discounts stack. Because batch requests run concurrently, though, cache hits are “best-effort”, with typical hit rates of 30% to 98%. It suggests the 1-hour cache TTL for batches with shared context. Google supports context caching inside batch requests. Its batch page and pricing page describe the cached-token rate inside a batch differently, so check both before you budget for it. OpenAI’s pricing page lists batch rates for cached input and cache writes on its newer models.
Both OpenAI’s and Google’s pricing pages also list a Flex tier, at the same rates as Batch for the models we checked. We haven’t covered how Flex behaves here. Read its own docs before relying on it.
How each batch API works
| Anthropic Message Batches | OpenAI Batch API | Gemini Batch API | |
|---|---|---|---|
| Input | Array of requests in the create call, each with custom_id and params |
.jsonl file uploaded via the Files API, one request per line |
Inline requests (under 20 MB total) or a JSONL file via the File API |
| Size limit | 100,000 requests or 256 MB, whichever comes first | 50,000 requests or 200 MB input file | 2 GB per input file |
| Window | Expires if not finished in 24 hours; most finish within 1 hour | completion_window must be 24h |
Target 24 hours; job expires after 48 hours pending or running |
| Result retention | 29 days after creation | Output file deleted 30 days after completion | 6 weeks by default |
| One model per batch? | Not stated; each request carries its own params |
Yes: one model per input file | Model is set on the batch job |
Some details that bite:
- Ordering. Anthropic and OpenAI both state that results may not come back in input order. Match on
custom_id. - Validation is late at Anthropic. Validation of each request’s
paramshappens asynchronously, and errors come back only when the whole batch ends. Test one request against the normal Messages API first. - Unsupported parameters. Anthropic rejects
stream: trueandmax_tokens: 0in batch requests. - Separate rate limits at OpenAI. Batch has its own pool with queued-token limits per model and up to 2,000 batch creations per hour. Batch usage does not consume your standard per-model limits.
- Expiry is still billed at OpenAI. If a batch expires, unfinished requests are cancelled, and you are charged for tokens consumed by requests that did complete.
- Gemini job creation is not idempotent. Sending the same create request twice makes two jobs and two bills. Google’s best-practices section says this directly.
- Spend limits can overshoot. Anthropic notes that batches can slightly exceed a workspace’s configured spend limit because of concurrent processing.
A minimal OpenAI-style input file looks like this (the model name is a placeholder for whichever model you use):
{"custom_id": "doc-000001", "method": "POST", "url": "/v1/responses", "body": {"model": "<model>", "input": "Classify this ticket: ..."}}
{"custom_id": "doc-000002", "method": "POST", "url": "/v1/responses", "body": {"model": "<model>", "input": "Classify this ticket: ..."}}
Gemini’s file format is similar, with a user-defined key and a request object per line. Anthropic takes the requests directly in the create call:
curl https://api.anthropic.com/v1/messages/batches \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"requests": [
{"custom_id": "doc-000001",
"params": {"model": "<model>", "max_tokens": 200,
"messages": [{"role": "user", "content": "Classify this ticket: ..."}]}}
]}'
Which workloads fit
A batch API fits when three things are true: nobody is waiting on the individual result, the work can be split into independent requests, and a delay of up to a day (or two, for Gemini’s expiry window) is acceptable.
Good fits, several of which the providers list themselves:
- Classifying or tagging large datasets (support tickets, product catalogs, documents)
- Extracting structured fields from documents
- Summarizing archives or nightly logs
- Running evaluations against a test set after a prompt or model change
- Generating embeddings for a content repository (OpenAI supports
/v1/embeddingsin batch) - Content moderation backlogs
- Bulk image generation (OpenAI and Google both document batch image generation)
Poor fits:
- Chat, autocomplete, or anything interactive
- Multi-step agent loops where step N depends on step N-1’s output. Each step would wait for a batch round trip. Anthropic does support server-side tool loops in batch and returns
pause_turnwhen a turn needs continuing, but the latency compounds. - Jobs with a hard deadline under 24 hours, unless you can absorb an occasional expiry and resubmit synchronously
A common pattern is a hybrid: serve users synchronously, and push everything that can wait (back-fills, re-scoring, evals, nightly digests) into batches.
Worked example: classifying 300,000 documents a month
This is an arithmetic example using published prices, not a measured bill. The assumptions:
- 300,000 product descriptions a month.
- Each request is about 1,200 input tokens (instructions plus the description) and 150 output tokens (a JSON label set).
- Monthly totals: 360M input tokens and 45M output tokens.
- Token counts are treated as equal across vendors for readability. They are not: Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text (more on that).
| Model | Standard price | Batch price | Saved per month |
|---|---|---|---|
| Claude Sonnet 5.5 ($2 / $10) | $1,170.00 | $585.00 | $585.00 |
| Claude Haiku 4.5 ($1 / $5) | $585.00 | $292.50 | $292.50 |
| Gemini 3.8 Flash ($0.75 / $3.75, through Dec 31, 2026) | $438.75 | $219.38 | $219.37 |
| Gemini 3.8 Flash from Jan 1, 2027 ($1.50 / $7.50) | $877.50 | $438.75 | $438.75 |
| GPT-5.6 Luna ($0.20 / $1.20) | $126.00 | $63.00 | $63.00 |
Two things stand out. Batch roughly matches what dropping a model tier would save, without changing models. And from January 2027, Gemini 3.8 Flash in batch costs exactly what it costs at standard rates today (details on that increase).
How many batches? Each limit applies to whichever is reached first, so check both request count and file size. As a rough estimate, a 1,200-token request is about 5 KB of JSONL once you include the JSON wrapper (around 4 characters per token plus overhead). At that size:
- OpenAI: 200 MB ÷ 5 KB ≈ 40,000 lines per file. That is below the 50,000-request cap, so plan on 8 files a month.
- Anthropic: 256 MB ÷ 5 KB ≈ 51,000 requests. That is below the 100,000 cap, so plan on 6 batches.
- Gemini: 2 GB ÷ 5 KB ≈ 400,000 lines. One file fits a month’s work, but splitting it gets you intermediate results sooner, which Google recommends for large jobs.
Measure your own line sizes; requests with long documents hit the byte limit much sooner. You can compare the monthly totals for other models in the LLM API cost calculator.
Checklist before moving a job to batch
- Confirm the latency budget. Can the consumer wait 24 hours, and handle the rare expiry (24 hours at Anthropic and OpenAI, 48 at Google)?
- Give every request a stable ID and write results keyed by it. Never rely on order.
- Validate one request synchronously first, especially at Anthropic, where batch validation errors arrive only at the end.
- Split by size and by model. Respect both the request and byte limits, and at OpenAI keep one model per file.
- Make submission idempotent yourself. Record the batch or job ID before retrying a create call. Gemini creates a duplicate job if you resend.
- Download results promptly. Anthropic keeps them 29 days, OpenAI deletes output files 30 days after completion, and Gemini keeps them 6 weeks.
- Put the stable part of each prompt first so any caching that does happen inside the batch can hit (how caching works).
- Cap the spend. A malformed batch of 100,000 requests is billed like 100,000 requests. Put a per-run cost cap in front of whatever generates the file.
Prices and limits change. Confirm them against the provider pages listed in the sources before you size a production job.