What AI scripts broke on our site: 8 defects with numbers

What AI scripts broke on our site: 8 defects with numbers

Last verified: September 29, 2026
12 min read
Case study
Technical SEO
AI integration

Each of the eight defects described below went through a script or an agent that finished with a success message. None of them threw an exception. All of them reached production on a multilingual site we have been building and maintaining for years, and we found every one only when we started comparing the output with something other than the tool’s own report.

In short: between August and September 2026, AI bulk passes put 22,202 filler paragraphs on our site, destroyed a heading, left 358 English sentences on pages in five other languages and produced 518 incorrect German compounds. Agents writing city pages added invented case studies and 30 links to pages that do not exist. We describe the mechanism behind each defect and the gate that catches it today.

We are writing this about our own site because only here do we have the full data: the commit, the file count, the state before and after. Search Engine Land published a piece this week on ten ways Claude can derail SEO if nobody checks its work. This is the version with receipts.

#Scale and context

wppoland.com runs on Astro and is built statically in six languages: Polish, English, German, Norwegian, Portuguese and Spanish. The last production build, on 29 September 2026, generated 13,733 pages. Content is written and revised by coding agents and by Node scripts that go through thousands of files at once. We have more than 70 quality gates in CI and locally.

Even so, every defect described here passed all of them. The reason is the same every time, and we come back to it at the end.

DefectWhat reached productionHow we found itGate today
Padding to a word count22,202 paragraphs in 873 files59 copies of one paragraph on a live pagecheck:no-batch-filler
City page filler7,567 blocks in 2,621 filesNumbered headings “Ateny (1)” to “Ateny (26)“the same one, extended to cities
Price regex1 destroyed headingGlued-words scannercheck:glued-words
Translating only conjunctions358 English sentences in 104 filesComparison with the English versionwpId-based detector
Term replacement without compounds518 compounds in 356 filesReading the page by handsearching for the class, not the literal
Invented case studies13 pages, 30 dead linksPre-publication review of the agents’ workinternal link validation
A deploy that re-warmed the old cache5 of 6 home pages with old contentComparing content, not HTTP codesa second purge after the smoke test
Deepening that replaced content142 to 206 elements per commitComparison with the previous version of the filecheck:element-loss

#1. Padding to a word count: 22,202 paragraphs

Our editorial guidelines say a blog post runs to about 2500 words. That was guidance for a human. A series of batch scripts treated it as a target to hit and padded shorter texts with generated filler.

The result, measured during the clean-up on 29 August: 22,202 numbered paragraphs such as “Notatka wdrożeniowa 47” (implementation note 47) in 873 files, up to 60 identical copies in a single post, plus 4,144 template sections rotated from a pool of five per language. Two of those sections were internal editorial instructions, published as article body. On the live page about WordPress security updates, the same paragraph stood 59 times.

The script had no bug. It did exactly what it was asked to do: raise the word count. The mistake was treating the word count, a supporting measure, as the goal.

Gate today: check:no-batch-filler in every npm run check run. The first version looked for the generator’s literals and let through filler that came back with different text. The current one looks at shape: the same second-level heading three times in one file is an error, whatever the words.

#2. City pages: the same filler, a different directory

For weeks the city pages collection slipped past every gate, because none of them scanned it. Passes 11 to 14 padded these pages to 2200 and then 2500 words. During the clean-up on 30 August we removed 7,567 numbered blocks and around 21 thousand repeated sections from 2,621 files. The page /pl/woocommerce-programista-ateny/ showed sections from “Ateny (1)” to “Ateny (26)”.

Worse, one of the passes appended blocks in Spanish to every language except Norwegian. Polish, German and Portuguese city pages carried “Entrega y seguimiento” sections.

The second lesson came during removal. The first version of the clean-up script compared whole literals and reported success while leaving 9,642 sections behind, because later repair scripts had edited that filler in place (for example a Spanish gender fix in 1,506 files). An exact byte comparison could not see any block a fixer had touched. We now match the heading plus the first four words.

After the filler was removed, 724 pages fell below the 2000-word threshold. We checked them in Google Search Console before deciding anything: only 2 had any click in 90 days, 575 had zero impressions. We moved 723 to noindex instead of padding them back up.

#3. The price regex that ate a heading

Service pages do not show prices outside the pricing page, so a script replaced amounts in zloty with “wycena indywidualna” (individual quote). “zł” is the Polish currency symbol, and the regex matched it case-insensitively with no word boundary after it, so it also matched the first two letters of any Polish word starting with “zł”.

On the Polish security audit page, the heading “2. Złośliwe przekierowania” (malicious redirects) looked to this regex like the price “2. Zł”. In production it read “wycena indywidualna ośliwe przekierowania”. The same regex sat in a second script, the one handling prices on city pages.

We found it by accident while building a scanner for headings with glued words, which on the same occasion found 13 headings missing a space before a preposition (“Is AMP deadin 2026?”). The fix is a negative lookahead in both scripts. Across the whole corpus one heading fell victim, but it was the heading of a sales page.

#4. Translation that only swapped conjunctions

The facts field (llmCard) is displayed visibly under articles and goes into the data for language models. On pages in the five languages other than English, 358 of these sentences were in English, in 104 files. Some were half translated: an old script had replaced only the conjunction, so a German portfolio page read “Contact und inquiry forms” and a Polish one “granite, conglomerate, i marble countertops”.

Detection was the interesting part. The first detector, based on the share of English function words, found 172 sentences. The translation agents themselves reported that other English sentences remained in the same files, and they were right. The second pass compared every sentence with the facts of the English version of the same page (the same wpId). Without an extra filter it returned 4,004 hits, mostly technology lists on city pages and the Norwegian “for”. With a requirement of at least two English function words, 170 real ones remained. We fixed the last 12 by hand.

We checked each of the five translations with a script, not with the agent’s report: in every file only the listed entries changed, every number survived, and there are no long dashes.

#5. A term replacement that did not know about compounds

A pass standardising vocabulary replaced a German term with “laufende Betreuung”. It did not handle compounds, so pages read “laufende Betreuung-Commitments”, “laufende Betreuung-Übergabe” and “Wartungs-laufende Betreuung”. That is not German. In total, 518 occurrences in 356 files, all on indexed pages.

The item in our backlog had a verification command that searched only for the form with a hyphen after the term. Fixing exactly what it pointed to would have turned it green and left 157 compounds in the second form in 118 files. So we searched for the defect class (any hyphen touching the term), not for the literal someone had typed into the task.

The repair method was simple: 497 of the 518 occurrences came from four template sentences. We rewrote each one once, by hand, in correct German (“Übergabe in die laufende Betreuung”, “Wartungsbetreuung”) and swapped it in as an exact sentence. The remaining 21 we fixed one by one.

#6. Agents that invent references

Thirteen curated city pages were written by agents. The pre-publication review on 26 August found sections in them such as “Case Study 1: Dystrybutor B2B z Bielan Wrocławskich” (a B2B distributor from Bielany Wrocławskie) with precise numbers: +52% enquiries, LCP from 4.5 s to 0.7 s, 100/100 in PageSpeed. None of these clients exists. The same numbers also sat in the facts field and in the speakable data, not only in the body.

On top of that, the “Other locations” sections linked to cities guessed on geography: Lübeck, Kiel, Regensburg, Girona, Tromsø. 30 dead links in 8 files.

No text gate catches this, because an invented case study is syntactically correct. A process rule catches it: the agent receives a list of existing URLs in its instructions, and every numerical claim about a client needs a source in the repository or it goes.

#7. A deploy that succeeded, production with old content

The deployment script on 8 September finished cleanly: upload, cache purge, smoke test 81/81 OK, exit code 0. On the live site, five of the six home pages showed the old tiles.

The order was: upload, purge, an 8-second pause, smoke test. Cloudflare Pages had not yet activated the new deployment, so the test’s 81 requests hit the previous build and filled the cache with it for an hour. The verification step undid the purge performed a moment earlier. No signal was false. None measured what went wrong: the test checked HTTP codes, not content.

Today: the script purges the cache a second time, after the test, and we confirm the deployment by finding a phrase from the specific change on the live site, with a cache-busting parameter.

#8. Deepening that replaced content

Passes “deepening” short posts were supposed to add content. In practice some of them replaced the entire article body. We restored nine posts from git history.

After that we built a gate that compares every changed file with its base version and reports the loss of an element: a table, a component, an iframe, an image, a code block, an FAQ question, a howTo step, an internal link. Run backwards over the last 60 content commits, it flagged every deepening pass from 128 to 185, each of which lost between 142 and 206 elements, including tables, embedded videos and code blocks.

The first version of this gate recognised headings by their text. On the current branch it reported 17 lost headings and all 17 were fixes: glued words split apart, an eaten heading restored. A rename is not a loss. So we count headings, links and FAQ entries by quantity, and keep identity only for images, components and iframes.

#The shared mechanism: success from the tool’s point of view

All eight cases have the same shape. The tool measured what it had done itself and reported success on that basis. The padding script measured the word count. The clean-up script measured whether it had found its literals. The backlog verification measured one form of the compound. The smoke test measured HTTP codes.

We found each defect only when we compared the output with something external:

  • with the previous version of the same file (lost tables and sections),
  • with the version in another language (English sentences on German pages),
  • with the live site instead of the deployment log (the old cache),
  • with the defect class instead of the literal from the task (the second form of the compounds),
  • with demand data (723 city pages without a single impression).

This gives the practical rule we have applied since September: an agent’s or a script’s report is a hypothesis. The result is checked by a separate script that does not know what the tool meant to do, and compares the state before with the state after.

#What I would change before the first bulk pass

  1. No numerical target for a script that writes prose. Word count, link count and FAQ count are measures to read, not to meet.
  2. A gate comparing against the previous version of the file before the first batch commit, not after the thirtieth.
  3. Every regex on text tested against headings and against words with Polish characters, because “Zł” is the start of many Polish words.
  4. Translations verified by comparison with the version in the source language, not with a word list.
  5. A task’s verification command searches for the defect class. If the task names one example, search for its variants too.
  6. A deployment confirmed by content on the live site, in every language.

#What this post does not prove

We do not claim these defects cost us traffic, because we have not measured it in a way that would settle the question. Some of the pages with filler had no impressions before it either.

Google’s September spam update started on 25 September 2026 and is expected to take about two weeks. According to the Search Engine Roundtable summary of 28 September, it hit programmatic pages and AI-generated content, and in the same week Google published a paper on the SAFE system for detecting mass “AI slop”. Our city pages are exactly that category. We will do the Search Console read, separately for city pages, the blog and service pages, once the update has finished, and we will write it up whatever the result.

One last note, about ourselves: this post was also written with the help of an agent. Every number in it comes from a commit or a measurement in our repository and went through the same gates we describe here.

Next step

Turn the article into an actual implementation

This block strengthens internal linking and gives readers the most relevant next move instead of leaving them at a dead end.

Want this implemented on your site?

If visibility in Google and AI systems matters, I can build the content architecture, FAQ, schema, and internal linking needed for SEO, GEO, and AEO.

Related cluster

Explore other WordPress services and knowledge base

Strengthen your business with professional technical support in key areas of the WordPress ecosystem.

Can AI damage a site's SEO when every step ends in success?#
Yes, and that is exactly what happened to us. A script padding texts to a word count, a script removing prices and a translation agent all finished without an error, and production received 22,202 filler paragraphs, a broken heading and 358 English sentences on pages in other languages. A tool reports success from its own point of view, not from the page's.
How do you detect English text left behind on translated pages?#
A word list only gets you part of the way. Our first detector, based on the share of English function words, found 172 of the 358 sentences. We found the rest by comparing every sentence with the English version of the same page (same wpId) and requiring at least two English function words, because similarity alone also caught lists of technology names.
Which gate catches content that a script silently removed?#
Comparing the file with its previous version. Our check:element-loss gate counts tables, components, iframes, images, code blocks, FAQ questions and howTo steps in the base version and in the new one, and reports every drop. Run backwards over 60 commits, it flagged every content deepening pass, each losing between 142 and 206 elements per commit.
Can mass-generated city pages stay in the index?#
Only the ones with their own content and evidence of demand. After the filler was removed, 724 city pages fell below our 2000-word threshold. Of those, only 2 had any click in 90 days and 575 had zero impressions, so we moved 723 to noindex instead of padding them back up.

Need an FAQ tailored to your industry and market? We can build one aligned with your business goals.

Let’s discuss

Related Articles

AI-slop content cleanup

A YMYL diagnostic for WordPress sites: how to find fake stats, fabricated citations, duplicate AI pages, wrong dates, and invented team bios before they damage trust, compliance, or AI citations.

Generative AI in Search Console

The new Generative AI section in Google Search Console reports impressions from AI Overviews and AI Mode. What it measures, what it leaves out, and how to read the data without drawing the wrong conclusions.