Korean Title Cards with Playwright Instead of AI Image Models
How we render Korean title, table and checklist cards from HTML with Playwright: keep-all wrapping, shrink-to-fit titles, font loading and element shots.
- Every image with words on it in our blog pipeline is an HTML template that headless Chromium screenshots. That covers title cards, table cards, checklist cards and section summary cards. Each costs $0 and shows exactly the text the post contains.
- The only paid image is one AI hero picture per post. Google lists a 1K Gemini 3.1 Flash Image output at $0.067, and it was 64–83% of each run’s total cost in our tests.
- Our first title card broke Korean words in the middle.
word-break: keep-allfixed it. - Titles use
white-space: pre, so the model controls line breaks. A small loop then shrinks the font until the longest line fits. - We wait for
document.fonts.readybefore the shot and capture the element, not the viewport. That’s why a 600 px “viewport” table card came out 1,145 px tall.
We built these cards for our client-facing blog, a small Korean web studio’s blog on Naver. The pipeline behind it is described in the architecture write-up. Korean blogs on that platform tend to be image-heavy. The research step in our pipeline found the top five posts for our test keyword carried between 3 and 27 images each. So every post needs several images, and most of them need readable Korean text.
Why HTML cards instead of an image model
We considered asking an image model for title cards and decided against it, for three design reasons. We didn’t run a head-to-head test, so this isn’t a benchmark.
- The text must match the post exactly, and our checks can’t read pixels. The title on a card comes from the same JSON that the SEO lint has already checked. If an image model renders that title, any misspelled syllable or dropped particle is baked into a PNG that no step downstream can verify. An HTML template can only show the string we hand it.
- Cost. Google’s pricing page lists Gemini 3.1 Flash Image output at $60 per million tokens, which it equates to $0.067 per 1K image. One post in our tests used 10 images. Seven were HTML renders and two were page screenshots, all free. Doing all ten with the image model would have cost about $0.67 for images alone. That’s several times our whole per-run cap. (The cap is explained here.)
- Consistency. Every card shares one stylesheet: the same fonts, paper colour, rules and footer. A reader scrolling the blog sees one visual system, not ten different moods.
What we kept the image model for is the one thing templates are bad at: a photographic hero image with no text. The prompt suffix in the code is explicit about that:
contents: [{ parts: [{ text: `${post.hero_prompt}. Clean editorial photo style, warm natural light, no text, no logos, no brand UI.` }] }],
generationConfig: { responseModalities: ['IMAGE'], imageConfig: { aspectRatio: '16:9', imageSize: '1K' } },
The hero is optional. If the call fails, the pipeline logs “hero image skipped” and carries on. The preview captions it as AI-generated, and that caption comes from the code, not the model.
The card templates
All templates share one BASE block: a Google Fonts link for Noto Sans KR and Noto Serif KR, plus about 20 lines of CSS. The card body is plain HTML built from the post JSON, with every value escaped. The title card, with the studio’s name and domain replaced by placeholders:
await shot(browser, `<div class="card"><div><div class="kicker"><STUDIO> · 제작 노트</div></div>
<div><h1 class="fit">${esc(post.card_title)}</h1><div class="sub">${esc(post.card_sub)}</div></div>
<div class="foot"><span><tagline></span><b><studio-domain></b></div></div>`,
path.join(dir, '01-title.png'));
The model supplies card_title and card_sub through the writing schema. The schema descriptions do the art direction: “2–3 line title for the card, line breaks (\n) at meaning boundaries, 12 characters per line max” and “card subtitle, 20 characters max”.
Table and checklist cards take the first section that has a table and the first that has a checklist. Sections with neither get a summary card: kicker, heading, and the first sentence of the section, up to 90 characters. We added summary cards after the research step showed how image-heavy the competing posts were. They raised our final test run to 10 images without any extra spend.
Fixing mid-word wrapping with keep-all
The first title card we rendered broke a Korean word in the middle of the line. That’s normal browser behaviour. By default, CJK text can wrap between any two characters, and to a Korean reader a word split across lines reads like a typo. MDN describes the value that changes this: with keep-all, word breaks “should not be used for Chinese/Japanese/Korean (CJK) text”, so lines break at spaces instead. The fix was one property on body:
body{width:1080px;word-break:keep-all;overflow-wrap:break-word;font-family:'Noto Sans KR',sans-serif;background:#f4efe6;color:#1b1a17}
overflow-wrap: break-word is the safety net. A single unbroken token longer than the line, such as a long product name or a URL, can still wrap instead of spilling out of the card. Because it sits on body, the rule covers table cells and checklist items too. In the table card from our final run, cells like a two-line product-category label wrap only at the space between words.
Shrink-to-fit for white-space: pre titles
For the big title we didn’t want the browser choosing line breaks at all. The model already breaks card_title into 2–3 meaningful lines. So the h1 uses white-space: pre, which keeps those newlines and never adds new ones. The catch is that a long line now overflows the card instead of wrapping. The fix is a loop that runs in the page after fonts load:
await page.evaluate(() => {
for (const h of document.querySelectorAll('h1.fit')) {
let fs = 92;
while (h.scrollWidth > h.clientWidth && fs > 40) { fs -= 2; h.style.fontSize = fs + 'px'; }
}
});
It starts at the design size of 92 px and steps down 2 px at a time until the content width fits the box, with a floor of 40 px so a runaway title can’t turn into fine print. Measuring scrollWidth against clientWidth works because pre content never wraps: any overflow is horizontal. The pipeline doesn’t record the final size, so we can’t say how often the loop fires. It’s there for the day the model ignores the 12-character guidance.
Wait for fonts, then shoot the element
Two details decide whether the screenshot looks like the design.
Fonts. The cards load Noto Sans KR and Noto Serif KR from Google Fonts. Korean font files are large, and if the screenshot fires before they arrive, Chromium renders a fallback system font. That font also has different widths, which would throw off the shrink loop. So the order is:
await page.setContent(`<!doctype html><html lang="ko"><head><meta charset="utf-8">${BASE}</head><body>${html}</body></html>`,
{ waitUntil: 'networkidle' });
await page.evaluate(() => document.fonts.ready);
// ...shrink-to-fit loop...
const el = await page.$('body > *');
await el.screenshot({ path: file, type: 'png' });
MDN says the document.fonts.ready promise resolves “once the document has completed loading fonts, layout operations are completed, and no further font loads are needed”. Returning it from page.evaluate makes Playwright wait for it. The fit loop runs after that, so it measures real glyph widths.
Element screenshots. We screenshot the card element (body > *), not the page. The viewport size is only a starting layout width, and the image takes the element’s real height. In our final run the table and checklist cards were opened with a 600 px viewport. The table card came out 1080×1,145 and the checklist card 1080×935, because their content was taller than 600 px. The fixed-height summary cards came out at exactly 600 px. For square title cards the element itself is 1080×1080, so the output is exact.
Real screenshots for “real screen” images
Not every image should be a designed card. Each post also links to one of the studio’s service pages and a past project. The pipeline screenshots both as they are live, at a 1280×860 desktop viewport:
await page.goto(u, { waitUntil: 'networkidle', timeout: 45000 });
await page.waitForTimeout(1500);
await page.screenshot({ path: path.join(dir, f) });
The extra 1.5 seconds lets animations and lazy-loaded hero images settle. This is a plain viewport shot, not full page: we want “this is what the page looks like”, not a 9,000 px strip. A page that fails to load is skipped with a log line, and the post keeps its other images. Since each post links to the service page that fits its topic, these screenshots also vary between posts. A blog that reuses identical images looks templated to readers.
One operational note: the first run on a fresh machine failed with browserType.launch: Executable doesn't exist at ~/Library/Caches/ms-playwright/.... Installing the npm package doesn’t download a browser. The fix was the command the error suggested:
npx playwright install chromium
What we’d do differently
- Self-host the fonts. Every card depends on Google Fonts loading within the
networkidlewindow. Bundling the two Noto families locally would remove a network dependency and make renders reproducible offline. - Use locators.
page.$('body > *')returns an element handle. Playwright’s docs show element screenshots throughpage.locator(...).screenshot(), which is the current style. - Log the final font size. If the shrink loop ever reaches its 40 px floor, the card is probably unreadable. That should be a warning in
run.json, not something we notice by eye. - Check alt text against the keyword list. Alt text is generated from the main keyword and section heading. It’s cheap to lint, and the platform’s image search reads it.
To see what the paid hero image does to a budget at different volumes, the LLM API cost calculator handles the text side. Add $0.067 per 1K image on top.