Gemini responseSchema in Production: What Structured Output Fixed

Lessons from using Gemini JSON mode with responseSchema: minItems for length, required fields, schema descriptions, source_ids validation, Korean keyword counting.

By the ailog editors · Published Oct 1, 2026 · 8 min read · How we work
In short
  • A draft that came out too short didn’t get longer after one prompted retry. It did once we added minItems and maxItems to the sections array and its paragraphs. The next run went from 1,449 to 2,862 characters.
  • Putting a field in required is how you get the model to tell you things. Our required unsupported_claims array lists every fact the model chose not to state.
  • description strings inside the schema work as per-field instructions, and Google’s guide recommends them.
  • Schema-valid isn’t the same as true. We validate each section’s source_ids against the real source list after parsing, as the docs advise.
  • For Korean SEO, count keywords with spacing ignored. The fix was correct but didn’t change our failing draft’s count. It used the keyword too few times, plain and simple.

These notes come from a Node.js pipeline we built that writes Korean posts for our client-facing blog, a small Korean web studio’s blog on Naver. The full architecture is here. Every model call in it uses Gemini’s JSON mode with a schema. We use gemini-3.8-flash for text, over the REST generateContent endpoint. Below is what the schema did for us, what it didn’t, and where the docs and our code differ.

The setup, and a note on field names

Our calls put the schema in generationConfig:

const { parts } = await gemini('gemini-3.8-flash', {
  contents: [{ parts: [{ text: prompt }] }],
  generationConfig: { responseMimeType: 'application/json', responseSchema: SCHEMA, temperature: 0.6 },
});
const post = JSON.parse(parts.map((p) => p.text || '').join(''));

When we checked Google’s structured-output guide on October 1, 2026, its examples had moved to a response_format object with mime_type and schema keys, shown against the newer Interactions endpoint. The generateContent API reference still documents responseMimeType, responseSchema and responseJsonSchema under GenerationConfig, and that’s what our code uses. Both work. Just don’t be surprised when a snippet from the guide doesn’t drop straight into a generateContent call.

The guide lists the schema keywords that are supported: enum, format, minimum, maximum, minItems, maxItems, items, prefixItems, properties, required and additionalProperties, plus title and description. It also warns that “not all JSON Schema features are supported” and that very large or deeply nested schemas may be rejected. Our writing schema is two levels deep with 13 required top-level fields. It never got rejected.

minItems fixed short posts where a retry didn’t

The target platform rewards substantial posts. Our lint wants the body between 1,500 and 3,600 characters. The first schema version had sections and paragraphs as plain arrays, and the prompt asked for 5–6 sections of 3–4 paragraphs.

The first lint-checked run came back with four sections and a 1,449-character body. The lint flagged it, and the pipeline made its one automatic rewrite, with the problem list added to the prompt. The rewrite was still too short. The run ended at 1,439 characters and failed the lint. So asking again in words didn’t help.

The fix was to move the numbers into the schema:

sections: {
  type: 'array',
  minItems: 5,
  maxItems: 6,
  items: {
    type: 'object',
    properties: {
      heading: { type: 'string' },
      paragraphs: { type: 'array', minItems: 3, maxItems: 5,
        items: { type: 'string', description: 'one paragraph, 120–220 characters' } },
      // table, checklist, source_ids ...
    },
    required: ['heading', 'paragraphs', 'source_ids'],
  },
},

The next run, a few minutes later with the same brief, came back with five sections and 2,862 characters, and passed every lint check on the first try. (Paragraphs were maxItems: 4 at that point. We raised it to 5 later, with a slightly longer paragraph hint.) Since then, every run has had 5 or 6 sections.

The general lesson: a count in the prompt is a suggestion, while a count in the schema shapes the output. Character length still can’t be set this way, because there’s no keyword for “this string is 120–220 characters”. The paragraph length stays a description hint. But fixing the number of containers does most of the work.

One caveat we found in our own code: schema limits apply to what the model returns, not to what your code does with it afterwards. Our post-filter drops sections with bad citations (below). That can bring a valid six-section reply under the minimum. The lint catches it as a length problem, but if you depend on counts downstream, check them after your own filters too.

required is how you get the model to admit things

Every top-level field in our writing schema is required. Most of that is plumbing: title, intro, sections, tags. The one that changed how we review drafts is this:

unsupported_claims: { type: 'array', items: { type: 'string' },
  description: 'content left out because there was no support for it (for review)' },

The prompt says facts may only come from the fetched SOURCES text, and that anything without support goes into this field. Because the field is required, the model can’t quietly skip the admission. It has to return an array, and in practice it filled it every time. The items it listed on our test runs were exactly the gaps a human reviewer would want to know about. The general card-fee rate wasn’t on the official fee page in a form it could use. Neither was the review lead time for signing up. A security incident that competing posts mentioned had no official source among ours. An optional field would most likely have come back empty.

Descriptions are instructions the model sees at the right moment

Google’s guide puts it directly: “Use the description field to guide the model.” We use descriptions as per-field instructions, and they’re where the most specific rules live:

Field Description (translated)
title 25–35 chars, starts with main_keyword (or has it within the first 10 chars), no clickbait
card_title 2–3 line title for the image card, break lines at meaning boundaries, 12 chars per line max
hero_prompt English prompt for the hero image; no text, logos or real brand UI
trending_issues (research schema) Recent issues seen in top posts; unverified, so don’t state them as fact before checking official sources
source_ids The SOURCES numbers this section’s facts are based on

Putting a rule next to the field it governs worked better for us than a long rules block at the top of the prompt. It’s also self-documenting: whoever reads the schema reads the editorial policy.

On enum: we didn’t use it, and we should have. The model picks a route_key (which service page the post links to) from a list we send in the prompt. The schema types it as a free string, and the code falls back to 'homepage' if the value isn’t a known key. Since enum is on the supported list, the stricter design is route_key: { type: 'string', enum: [...keys] }, which would remove the fallback. That’s advice from hindsight, not something we measured.

Validate source_ids against reality

A schema can force a field to be an array of integers. It can’t force those integers to point at anything. Google’s guide says as much: “While output is syntactically correct JSON, always validate values in your application.” So every section must cite at least one source number that actually exists:

post.sources = sources.map((s, i) => ({ id: i + 1, url: s.url }));
const valid = new Set(post.sources.map((s) => s.id));
post.sections = post.sections.filter((s) =>
  s.source_ids?.length && s.source_ids.every((id) => valid.has(id))
  || /<STUDIO>|체크|정리|마무리/.test(s.heading));

The numbering is ours. Sources are numbered in the prompt as [1], [2], [3] after fetching, and only pages that returned HTTP 200 with more than 200 characters of text get a number. A section that cites [4] when only three sources loaded is dropped. The regex exception lets through a closing tips or checklist section whose heading names the studio. That’s deliberate, but it’s also a hole: those sections are checked against the brand file by the prompt only. If you copy this pattern, flag exempt sections in the output rather than passing them silently.

This doesn’t prove a sentence is true to its source. It proves the model at least claimed a real source for each section, and the reviewer can follow that number to a URL.

Counting Korean keywords with spacing ignored

Korean spacing in compound nouns varies. A brand name plus “integration” can be written with or without a space, and readers treat both forms as the same search. Our lint originally counted keyword hits with an exact-match regex. After a run where the chosen main keyword itself contained a space, we switched to a counter that allows optional whitespace between every character:

const kNo = k.replace(/\s/g, '');
const re = new RegExp(kNo.split('').map((c) => c.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')).join('\\s*'), 'g');
const count = (txt) => (txt.match(re) || []).length;

Title placement is checked the same way: strip all spaces from title and keyword, then make sure the keyword starts within the first 10 characters.

To be honest about the result, we re-ran both counters on the saved draft that had failed the keyword check. Both found 3 hits. That draft simply used the spaced form three times and the unspaced form never, so the failure was a real shortfall, not a spacing artifact. The next run used the keyword four times and passed. We kept the new counter because it’s correct for Korean, but it didn’t rescue that draft. (A side effect to watch: \s* between every character can match across a line break or a word boundary, so very short keywords could overcount.)

Checklist

  • Put structural counts (minItems, maxItems) in the schema, not just the prompt.
  • Make “what did you leave out” a required field.
  • Write field-level rules as description strings.
  • Use enum for any value that must come from a fixed list.
  • Validate cross-references (source numbers, IDs, URLs) in code after parsing, and re-check counts after your own filters.
  • Normalise language-specific quirks, such as Korean spacing, before counting anything.
  • Check the docs for field names. The guide and the API reference currently show different shapes.

What these calls cost us per post, and the cap that stops a bad run, are covered in the per-run cost cap article. For your own token mix, try the LLM API cost calculator.

Sources
  1. Gemini API: Structured outputs
  2. Gemini API reference: models.generateContent (GenerationConfig)

Related