Ghost Duplicate Content Cleanup: 44 Posts Fixed via the Admin API

How we found every article on our Ghost blog stored twice (one post nested 196 copies), fixed it via the Admin API, and unpublished 19 posts.

By the ailog editors · Published Oct 1, 2026 · 9 min read · How we work
In short
  • ailog.page was flagged in AdSense as “Needs attention – Low value content”. When we looked inside the Ghost database, 44 of the 76 published posts had problems, and in 43 of them the whole article was stored twice: an old full-article HTML card plus the native Lexical body.
  • The worst post contained 14 copies of that HTML card, and each card contained 14 nested copies of itself. That’s 196 repeats of one section, about 4.7 million characters of stored HTML for a single article.
  • We fixed it through the Ghost Admin API: keep one copy, keep the “Key Takeaways” box, drop empty paragraphs, then PUT the new Lexical JSON with updated_at. Every public page was afterwards under a 5% repeated-sentence ratio, and between about 10,800 and 29,800 characters long.
  • We also unpublished 19 posts whose titles promised first-person tests that never took place. That content was the bigger problem, and ended up being the reason we rebuilt the site.

The rejection and what we expected to find

The AdSense site list showed ailog.page as “Needs attention”, reason “Low value content”, last updated 1 June 2026. The site was a self-hosted Ghost blog with 76 published posts, most of them produced by an automated publishing pipeline in early 2026. The ads.txt file and the ad code were already correct, so the problem was the content, not the setup.

We expected the usual low-value suspects: thin posts and generic rewrites. What we found was worse, and a lot easier to measure. A first crawl of the sitemap, counting only the pages that loaded, reported a median of about 4,100 words per post and a maximum of about 20,500. Real articles on these topics run 1,500–2,500 words. That alone told us the stored pages were much bigger than what anyone had written.

Diagnosis: from page sizes to node-type dumps

More than half the public pages failed to load when we crawled them, so we stopped crawling and queried Ghost’s MySQL database directly. Before touching anything, we took a dump of the posts tables (posts, posts_meta, posts_tags, post_revisions, mobiledoc_revisions), about 9.4 MB gzipped. Then we ranked posts by stored size and counted a heading that should appear once per article (“Real AI Responses”):

html_chars  lexical_chars  heading_count  slug
4,677,585   4,771,558      196            your-ai-writes-the-code-now-another-ai-reviews-it
  892,250     932,894       22            (next largest)
  542,599     571,934       18
  ...
21 posts over 60,000 characters of plain text

A heading appearing 196 times in one article isn’t a writing problem. Something was copying the content. To find out what, we dumped each post’s top-level Lexical nodes with their type, JSON length and first few words:

// inspect.mjs: one line per top-level Lexical node
const ch = JSON.parse(post.lexical).root.children;
const txt = (n) => n.type === 'html'
  ? n.html.replace(/<[^>]+>/g, ' ')
  : (JSON.stringify(n).match(/"text":"([^"]*)"/g) || []).map((s) => s.slice(8, -1)).join(' ');
ch.forEach((n, i) => {
  const t = txt(n).replace(/\s+/g, ' ').trim();
  if (n.type === 'html' || t) console.log(i, n.type, JSON.stringify(n).length, t.length, '|', t.slice(0, 70));
});

The pattern was the same everywhere. Each affected post had a large html node (a Ghost HTML card) holding a complete older version of the article, starting with a styled “Key Takeaways” box. After it came the same article again as native paragraphs, headings and lists. Visitors saw everything twice. An overlap script confirmed it: for each post with both a card and a native body, it checked whether each native paragraph also appeared inside the card. In 43 posts the answer was all of them, or all but one.

The biggest post had a second layer. It had 14 card nodes, and selfrep.mjs showed that each card’s HTML contained 14 back-to-back copies of the article. The copies were slightly different lengths, so they weren’t byte-identical pastes. 14 × 14 = 196, the number from the heading count. Our best guess is that an earlier publishing script fetched each post’s rendered HTML and wrote it back as a card on top of the existing content, again and again. We didn’t confirm that, and it doesn’t change the fix.

To have a number we could check before and after, we defined a duplicate ratio: split the text into sentences, keep those of 60 characters or more, and count the share that appear more than once.

function dupRatio(text) {
  const s = text.split(/(?<=[.!?])\s+/).filter((x) => x.length >= 60);
  const m = new Map(); s.forEach((x) => m.set(x, (m.get(x) || 0) + 1));
  return s.length ? s.filter((x) => m.get(x) > 1).length / s.length : 0;
}

Before the fix, most affected posts scored between 93% and 100%.

The fix: one copy per post, through the Admin API

Our first script was too simple. It hashed each top-level node and removed exact repeats. In a dry run it touched only 21 posts, and even in those it kept one full card plus the native body, so the article would still be on the page twice. We threw it away and wrote fix.mjs, which decides per post which single copy to keep:

  • Mode A: the native body is substantial (5,000+ characters of text). Drop every large HTML card (over 5,000 characters of HTML). If the card’s “Key Takeaways” box isn’t already in the native body, extract it and put it back at the top as a small HTML card. That box was the only useful thing the old card had that the native version didn’t.
  • Mode B: there’s little or no native body. Keep exactly one large card, and drop any native paragraph whose text already appears in that card.
  • Both modes: drop empty paragraphs at the start and end, and collapse runs of them. The duplicated sections had been separated by stacks of empty paragraphs.

Authentication is Ghost’s documented Admin API scheme: split the key into id and secret, sign a short-lived HS256 JWT with the hex-decoded secret, and send it as Authorization: Ghost <token>. The update must include the post’s current updated_at. Ghost uses it for collision detection, so the script won’t silently overwrite an edit someone made in the meantime. Here’s the core of it, cleaned up and without credentials:

// fix.mjs (core). Usage: GHOST_ADMIN_KEY=<id>:<secret> node fix.mjs [--apply] [--only <slug>]
import crypto from 'node:crypto';
const HOST = 'https://ailog.page';
const APPLY = process.argv.includes('--apply');
const [kid, secret] = process.env.GHOST_ADMIN_KEY.split(':');

function token() {
  const b64 = (o) => Buffer.from(JSON.stringify(o)).toString('base64url');
  const now = Math.floor(Date.now() / 1000);
  const h = b64({ alg: 'HS256', typ: 'JWT', kid });
  const p = b64({ iat: now, exp: now + 300, aud: '/admin/' });     // 5 minutes max
  const sig = crypto.createHmac('sha256', Buffer.from(secret, 'hex')).update(`${h}.${p}`).digest('base64url');
  return `${h}.${p}.${sig}`;
}
async function api(method, path, body) {
  const r = await fetch(`${HOST}/ghost/api/admin${path}`, {
    method, body: body && JSON.stringify(body),
    headers: { Authorization: `Ghost ${token()}`, 'Content-Type': 'application/json', 'Accept-Version': 'v5.0' },
  });
  const j = await r.json(); if (!r.ok) throw new Error(`${r.status} ${JSON.stringify(j).slice(0, 300)}`); return j;
}

const isCard = (n) => n.type === 'html' && n.html.length > 5000;
// nt(n): plain text of a node; isEmpty(n): paragraph with no text; textOf(nodes): joined, normalised text

const { posts } = await api('GET', '/posts/?limit=all&formats=lexical&filter=status:published&fields=id,slug,updated_at,lexical');
for (const p of posts) {
  const lx = JSON.parse(p.lexical); const ch = lx.root.children;
  const cards = ch.filter(isCard); if (!cards.length) continue;
  const native = ch.filter((x) => !isCard(x));
  let out;
  if (textOf(native).length >= 5000) {                                      // mode A
    out = native;
    const kt = cards[0].html.match(/<div[^>]*>\s*<h2[^>]*>\s*Key Takeaways[\s\S]*?<\/ul>\s*<\/div>/);
    if (kt && !textOf(native).includes(strip(kt[0]).slice(20, 80)))
      out = [{ type: 'html', version: 1, html: kt[0] }, ...out];
  } else {                                                                  // mode B
    const cardText = strip(cards[0].html); let kept = false;
    out = ch.filter((x) => isCard(x) ? (!kept && (kept = true))
                                     : !(nt(x).length > 20 && cardText.includes(nt(x).slice(0, 60))));
  }
  out = out.filter((x, i, a) => !(isEmpty(x) && (i === 0 || isEmpty(a[i - 1]))));
  while (out.length && isEmpty(out.at(-1))) out.pop();
  if (JSON.stringify(out) === JSON.stringify(ch)) continue;
  console.log(p.slug, (dupRatio(textOf(ch)) * 100).toFixed(0) + '%', '->', (dupRatio(textOf(out)) * 100).toFixed(0) + '%');
  if (APPLY) { lx.root.children = out;
    await api('PUT', `/posts/${p.id}/`, { posts: [{ lexical: JSON.stringify(lx), updated_at: p.updated_at }] }); }
}

We ran it in three passes. First a dry run over everything, reading the before/after sizes and duplicate ratios. Then --apply --only <one-slug>, after which we opened the live page and checked “Key Takeaways” appeared exactly once. Then --apply on everything. It reported changed=44, and a second dry run reported changed=0. Some representative lines from the dry run:

A  3280k -> 13.6k  dup 100% -> 2%  nodes 67 -> 54   (the 196-copy post)
A   505k -> 17.5k  dup 100% -> 0%  nodes 82 -> 61
A    87k -> 15.4k  dup  99% -> 0%  nodes 105 -> 101
B    40k -> 20.2k  dup  98% -> 0%  nodes 3 -> 1

Sizes here are normalised text characters, not HTML. The worst post went from about 3.3 million characters of text to about 13,600.

Verifying the public pages, not just the database

A clean database doesn’t guarantee a clean page. Themes, cached HTML and code injection can all add content. So verify.mjs fetched every published post’s public URL with a cache-busting query string, pulled out the article body, and checked three things: HTTP 200, under 40,000 characters of text, and a duplicate ratio under 5%.

200  29834  0%  (largest page)
200  26213  0%
200  25584  0%
...
posts 76  bad 0  minChars 10780

All 76 passed. The eight largest pages showed 0% duplicates, and lengths ranged from 10,780 to 29,834 characters. In the fix log, the 196-copy post ended at 2%, still well under the threshold and nowhere near a duplicated section.

The part that wasn’t a bug: posts we couldn’t stand behind

Fixing duplicates made the pages readable, and that made the next problem obvious. A batch of titles promised first-person experiments that, as far as we could establish, never happened. These were claims like testing a tool for hundreds of hours, tracking a month of API bills, or a teacher’s prep time dropping by a percentage. The Key Takeaways boxes repeated them with specific numbers we couldn’t trace to any record. A reader or a reviewer has no way to tell an invented test from a real one, and that’s exactly the kind of content that deserves a “low value” label.

We didn’t edit the claims down. We unpublished those posts by matching titles against a pattern:

// draft-fabricated.mjs: set matching posts to draft (revert: status 'published')
const RX = /\bI (Tested|Tracked|Built|Watched|Spent|Used|Asked|Coded|Fed|Cloned|Replaced)\b|\bI'm a Teacher\b|Save Me \d+ Hours|for \d+ Hours|Extensions I Tested/i;
const hit = posts.filter((p) => RX.test(p.title));
for (const p of hit) await api('PUT', `/posts/${p.id}/`, { posts: [{ status: 'draft', updated_at: p.updated_at }] });
// total=76 drafted=19 remain=57

A plain SQL LIKE search had found 16. The regex found 19, because it also catches titles where the claim comes mid-sentence. Drafting is reversible. The posts stay in Ghost, and their public URLs now return 404. We checked one by hand afterwards, along with the home page (200).

That decision led to the version of the site you’re reading. Rather than rescue the remaining 57 posts one at a time, we rebuilt ailog.page with a written rule: first-hand sections describe only work that actually happened, from real logs and files. The cleanup scripts quoted above are an example of that.

What we’d do differently

  • Measure before guessing. A word count per post and one heading count found the problem in minutes. “Low value content” sounds like a writing-quality issue, but ours was structural first.
  • Back up the tables before any bulk API write. It costs seconds, and the dump is the only undo for 44 posts.
  • Dry-run, apply to one, check the live page, then apply to all. Our first dedupe script would have “fixed” 21 posts and left all of them still duplicated.
  • Always send updated_at. It turns a bulk script into something that can’t overwrite a human edit.
  • Check that publishing scripts are idempotent. Whatever added cards to these posts did it again on every run. Anything that writes content back to a CMS should be safe to run twice.
  • Never publish a test you didn’t run. No script fixes that afterwards. The only fix is to remove it.
Sources
  1. Ghost Admin API — Token authentication
  2. Ghost Admin API — Updating a post (updated_at for collision detection)
  3. Google AdSense Help — Eligibility requirements for AdSense

Related