Each of the eight defects described below went through a script or an agent that finished with a success message. None of them threw an exception. All of them reached production on a multilingual site we have been building and maintaining for years, and we found every one only when we started comparing the output with something other than the tool’s own report.
In short: between August and September 2026, AI bulk passes put 22,202 filler paragraphs on our site, destroyed a heading, left 358 English sentences on pages in five other languages and produced 518 incorrect German compounds. Agents writing city pages added invented case studies and 30 links to pages that do not exist. We describe the mechanism behind each defect and the gate that catches it today.
We are writing this about our own site because only here do we have the full data: the commit, the file count, the state before and after. Search Engine Land published a piece this week on ten ways Claude can derail SEO if nobody checks its work. This is the version with receipts.
Scale and context
wppoland.com runs on Astro and is built statically in six languages: Polish, English, German, Norwegian, Portuguese and Spanish. The last production build, on 29 September 2026, generated 13,733 pages. Content is written and revised by coding agents and by Node scripts that go through thousands of files at once. We have more than 70 quality gates in CI and locally.
Even so, every defect described here passed all of them. The reason is the same every time, and we come back to it at the end.
| Defect | What reached production | How we found it | Gate today |
|---|---|---|---|
| Padding to a word count | 22,202 paragraphs in 873 files | 59 copies of one paragraph on a live page | check:no-batch-filler |
| City page filler | 7,567 blocks in 2,621 files | Numbered headings “Ateny (1)” to “Ateny (26)“ | the same one, extended to cities |
| Price regex | 1 destroyed heading | Glued-words scanner | check:glued-words |
| Translating only conjunctions | 358 English sentences in 104 files | Comparison with the English version | wpId-based detector |
| Term replacement without compounds | 518 compounds in 356 files | Reading the page by hand | searching for the class, not the literal |
| Invented case studies | 13 pages, 30 dead links | Pre-publication review of the agents’ work | internal link validation |
| A deploy that re-warmed the old cache | 5 of 6 home pages with old content | Comparing content, not HTTP codes | a second purge after the smoke test |
| Deepening that replaced content | 142 to 206 elements per commit | Comparison with the previous version of the file | check:element-loss |
1. Padding to a word count: 22,202 paragraphs
Our editorial guidelines say a blog post runs to about 2500 words. That was guidance for a human. A series of batch scripts treated it as a target to hit and padded shorter texts with generated filler.
The result, measured during the clean-up on 29 August: 22,202 numbered paragraphs such as “Notatka wdrożeniowa 47” (implementation note 47) in 873 files, up to 60 identical copies in a single post, plus 4,144 template sections rotated from a pool of five per language. Two of those sections were internal editorial instructions, published as article body. On the live page about WordPress security updates, the same paragraph stood 59 times.
The script had no bug. It did exactly what it was asked to do: raise the word count. The mistake was treating the word count, a supporting measure, as the goal.
Gate today: check:no-batch-filler in every npm run check run. The first version looked for the generator’s literals and let through filler that came back with different text. The current one looks at shape: the same second-level heading three times in one file is an error, whatever the words.
2. City pages: the same filler, a different directory
For weeks the city pages collection slipped past every gate, because none of them scanned it. Passes 11 to 14 padded these pages to 2200 and then 2500 words. During the clean-up on 30 August we removed 7,567 numbered blocks and around 21 thousand repeated sections from 2,621 files. The page /pl/woocommerce-programista-ateny/ showed sections from “Ateny (1)” to “Ateny (26)”.
Worse, one of the passes appended blocks in Spanish to every language except Norwegian. Polish, German and Portuguese city pages carried “Entrega y seguimiento” sections.
The second lesson came during removal. The first version of the clean-up script compared whole literals and reported success while leaving 9,642 sections behind, because later repair scripts had edited that filler in place (for example a Spanish gender fix in 1,506 files). An exact byte comparison could not see any block a fixer had touched. We now match the heading plus the first four words.
After the filler was removed, 724 pages fell below the 2000-word threshold. We checked them in Google Search Console before deciding anything: only 2 had any click in 90 days, 575 had zero impressions. We moved 723 to noindex instead of padding them back up.
3. The price regex that ate a heading
Service pages do not show prices outside the pricing page, so a script replaced amounts in zloty with “wycena indywidualna” (individual quote). “zł” is the Polish currency symbol, and the regex matched it case-insensitively with no word boundary after it, so it also matched the first two letters of any Polish word starting with “zł”.
On the Polish security audit page, the heading “2. Złośliwe przekierowania” (malicious redirects) looked to this regex like the price “2. Zł”. In production it read “wycena indywidualna ośliwe przekierowania”. The same regex sat in a second script, the one handling prices on city pages.
We found it by accident while building a scanner for headings with glued words, which on the same occasion found 13 headings missing a space before a preposition (“Is AMP deadin 2026?”). The fix is a negative lookahead in both scripts. Across the whole corpus one heading fell victim, but it was the heading of a sales page.
4. Translation that only swapped conjunctions
The facts field (llmCard) is displayed visibly under articles and goes into the data for language models. On pages in the five languages other than English, 358 of these sentences were in English, in 104 files. Some were half translated: an old script had replaced only the conjunction, so a German portfolio page read “Contact und inquiry forms” and a Polish one “granite, conglomerate, i marble countertops”.
Detection was the interesting part. The first detector, based on the share of English function words, found 172 sentences. The translation agents themselves reported that other English sentences remained in the same files, and they were right. The second pass compared every sentence with the facts of the English version of the same page (the same wpId). Without an extra filter it returned 4,004 hits, mostly technology lists on city pages and the Norwegian “for”. With a requirement of at least two English function words, 170 real ones remained. We fixed the last 12 by hand.
We checked each of the five translations with a script, not with the agent’s report: in every file only the listed entries changed, every number survived, and there are no long dashes.
5. A term replacement that did not know about compounds
A pass standardising vocabulary replaced a German term with “laufende Betreuung”. It did not handle compounds, so pages read “laufende Betreuung-Commitments”, “laufende Betreuung-Übergabe” and “Wartungs-laufende Betreuung”. That is not German. In total, 518 occurrences in 356 files, all on indexed pages.
The item in our backlog had a verification command that searched only for the form with a hyphen after the term. Fixing exactly what it pointed to would have turned it green and left 157 compounds in the second form in 118 files. So we searched for the defect class (any hyphen touching the term), not for the literal someone had typed into the task.
The repair method was simple: 497 of the 518 occurrences came from four template sentences. We rewrote each one once, by hand, in correct German (“Übergabe in die laufende Betreuung”, “Wartungsbetreuung”) and swapped it in as an exact sentence. The remaining 21 we fixed one by one.
6. Agents that invent references
Thirteen curated city pages were written by agents. The pre-publication review on 26 August found sections in them such as “Case Study 1: Dystrybutor B2B z Bielan Wrocławskich” (a B2B distributor from Bielany Wrocławskie) with precise numbers: +52% enquiries, LCP from 4.5 s to 0.7 s, 100/100 in PageSpeed. None of these clients exists. The same numbers also sat in the facts field and in the speakable data, not only in the body.
On top of that, the “Other locations” sections linked to cities guessed on geography: Lübeck, Kiel, Regensburg, Girona, Tromsø. 30 dead links in 8 files.
No text gate catches this, because an invented case study is syntactically correct. A process rule catches it: the agent receives a list of existing URLs in its instructions, and every numerical claim about a client needs a source in the repository or it goes.
7. A deploy that succeeded, production with old content
The deployment script on 8 September finished cleanly: upload, cache purge, smoke test 81/81 OK, exit code 0. On the live site, five of the six home pages showed the old tiles.
The order was: upload, purge, an 8-second pause, smoke test. Cloudflare Pages had not yet activated the new deployment, so the test’s 81 requests hit the previous build and filled the cache with it for an hour. The verification step undid the purge performed a moment earlier. No signal was false. None measured what went wrong: the test checked HTTP codes, not content.
Today: the script purges the cache a second time, after the test, and we confirm the deployment by finding a phrase from the specific change on the live site, with a cache-busting parameter.
8. Deepening that replaced content
Passes “deepening” short posts were supposed to add content. In practice some of them replaced the entire article body. We restored nine posts from git history.
After that we built a gate that compares every changed file with its base version and reports the loss of an element: a table, a component, an iframe, an image, a code block, an FAQ question, a howTo step, an internal link. Run backwards over the last 60 content commits, it flagged every deepening pass from 128 to 185, each of which lost between 142 and 206 elements, including tables, embedded videos and code blocks.
The first version of this gate recognised headings by their text. On the current branch it reported 17 lost headings and all 17 were fixes: glued words split apart, an eaten heading restored. A rename is not a loss. So we count headings, links and FAQ entries by quantity, and keep identity only for images, components and iframes.
The shared mechanism: success from the tool’s point of view
All eight cases have the same shape. The tool measured what it had done itself and reported success on that basis. The padding script measured the word count. The clean-up script measured whether it had found its literals. The backlog verification measured one form of the compound. The smoke test measured HTTP codes.
We found each defect only when we compared the output with something external:
- with the previous version of the same file (lost tables and sections),
- with the version in another language (English sentences on German pages),
- with the live site instead of the deployment log (the old cache),
- with the defect class instead of the literal from the task (the second form of the compounds),
- with demand data (723 city pages without a single impression).
This gives the practical rule we have applied since September: an agent’s or a script’s report is a hypothesis. The result is checked by a separate script that does not know what the tool meant to do, and compares the state before with the state after.
What I would change before the first bulk pass
- No numerical target for a script that writes prose. Word count, link count and FAQ count are measures to read, not to meet.
- A gate comparing against the previous version of the file before the first batch commit, not after the thirtieth.
- Every regex on text tested against headings and against words with Polish characters, because “Zł” is the start of many Polish words.
- Translations verified by comparison with the version in the source language, not with a word list.
- A task’s verification command searches for the defect class. If the task names one example, search for its variants too.
- A deployment confirmed by content on the live site, in every language.
What this post does not prove
We do not claim these defects cost us traffic, because we have not measured it in a way that would settle the question. Some of the pages with filler had no impressions before it either.
Google’s September spam update started on 25 September 2026 and is expected to take about two weeks. According to the Search Engine Roundtable summary of 28 September, it hit programmatic pages and AI-generated content, and in the same week Google published a paper on the SAFE system for detecting mass “AI slop”. Our city pages are exactly that category. We will do the Search Console read, separately for city pages, the blog and service pages, once the update has finished, and we will write it up whatever the result.
One last note, about ourselves: this post was also written with the help of an agent. Every number in it comes from a commit or a measurement in our repository and went through the same gates we describe here.






