Batch API 50% Discount: When to Use It at Anthropic, OpenAI and Google
How the Claude, OpenAI and Gemini batch APIs work: the 50% discount, 24-hour windows, size limits, request formats, good-fit jobs and a monthly example.
What AI APIs and tools actually cost, with official prices and our own measured bills.
How the Claude, OpenAI and Gemini batch APIs work: the 50% discount, 24-hour windows, size limits, request formats, good-fit jobs and a monthly example.
A practical framework for picking an LLM API tier by cost: worked examples for a classifier, RAG chatbot and long-form writer, plus routing and escalation.
Gemini 3.8 Flash goes from $0.75/$3.75 to $1.50/$7.50 per 1M tokens on Jan 1, 2027. Worked bill examples, alternatives, and what to do before January.
How our blog pipeline prices every Gemini response from usage metadata, stops at $0.20 per run, and what five real runs cost (about $0.09–$0.11 each).
How prompt and context caching works at Anthropic, OpenAI and Google: minimum sizes, TTLs, write premiums, storage fees, and worked cost examples.