XML sitemap: five mistakes we found in our own

XML sitemap: five mistakes we found in our own

Last verified: September 29, 2026
10 min read
Case study
Technical SEO
500+ WP projects

On 27 September 2026 Joost de Valk published a piece arguing that almost nobody gets XML sitemaps right. He ran six sites, from Adobe to GOV.UK, through his Sitemap Inspector plugin and found the same things in every one: noindexed URLs, redirects, 404s, canonicals pointing elsewhere, one date across hundreds of entries. His thesis: “A sitemap should list the URLs a site wants indexed, and nothing else.”

We agree with the thesis and with the list. We used it on our own site. Between August and October 2026 we found five bugs in the wppoland.com sitemap generator, each one measured in production. Joost’s list would have caught three of them. No inspector that reads sitemap entries will catch the other two, because they were about what the sitemap did not contain. This piece covers both kinds.

Technical context: wppoland.com is an Astro site in six languages, with about 6700 URLs across its sitemaps. The sitemaps come from a custom integration, astro-sitemaps.mjs, not a plugin. That matters for the conclusions: every bug below lived in the code that reads data and builds the list, not in the list itself.

#Bug 1: build date as lastmod

On 24 August 2026, five different sitemaps had exactly one lastmod value for every entry, the same one, today’s. The integration computed new Date() once and stamped it on 3079 URLs. Every deployment told Google that everything had changed.

Google states the condition directly: it uses <lastmod> if the value is “consistently and verifiably” accurate to the page’s last change (Google Search Central, the page on building sitemaps, updated 8 July 2026). A date that jumps with every build fails that test, so Google stops trusting it for the whole domain, including where it would have been accurate. In samples from that period, the last crawl of some pages went back anywhere from three weeks to three months.

The fix substituted updatedDate from frontmatter, falling back to pubDate. Effect on the artifact: the Polish blog sitemap went from 1 to 98 distinct dates.

This fix invites a second-order mistake. It was tempting to use the lastVerified field for city pages, since every one of the 4733 files has it. Except 4697 of the 4733 files hold the same value in it. That is a bulk stamp, not a signal, and it would have moved the same defect from “today” to a different constant date. Before any date field goes into lastmod, count how many distinct values it has.

The August fix covered the blog, the portfolio and pages from the content collection. It did not cover hand-written .astro routes. Joost’s article prompted us to check again: on 29 September, in /pl/sitemap-pages.xml, 62 of 113 URLs still had the build date as lastmod, because no content date exists for them and the code then fell back to today. The fix (commit af97c7a887) takes the date of the source file’s last commit for such routes, from a single git log call per build, and when there is none, omits lastmod. The shared template [lang]/[slug].astro deliberately does not count as a source, because its date says nothing about any particular page. The same file after the fix: 87 of 113 URLs have lastmod, 42 distinct dates, the most common shared by 7 URLs, and 26 URLs have no date at all. In the Polish city pages sitemap, 684 of 692 entries carried the build date before the fix; after it, 8 entries have a real date: the city template has no date that could honestly be used, so the rest declare none. The fix has been in production since 29 September 2026, and the live file shows the same figures.

#Bug 2: the noindex filter guarded one field

After 723 city pages were demoted to noindex, Search Console kept reporting them as discovered. The pages themselves had a correct noindex. They were getting back to Google another way: sitemap-locations.xml listed them as language alternates (xhtml:link, including x-default) of pages that stayed in the index.

The integration’s noindex filter checked only <loc>. The alternates come from the page head, which lists every language version, so the filter never saw them. The fix keeps only the alternates whose URL is itself a <loc> in the sitemap. After it, the build produced 6407 URLs and 0 bad alternates.

The lesson goes beyond sitemaps: a filter on one field of an entry does not cover the other fields of that entry that carry a URL. A sitemap has several such places: loc, alternates, x-default, image:loc.

#Bug 3: 500 live pages left out of the sitemap

On 21 September 2026 the integration started skipping every URL matching the prefixes that the middleware in functions/_middleware.ts redirects. But the middleware redirects them only when the resource returns 404, and every sitemap candidate is a built page, so it never returns 404. The filter copied the rule without its condition.

Result: 500 live pages with index, follow dropped out of the sitemap, including the target of one of the redirects (/pl/audyt-bezpieczenstwa-wordpress/ matched the prefix audyt-bezpieczenstwa-). The fix landed on 26 September, and the build with the sitemap gave +500 URLs and 0 removed.

No entry check will catch this bug. Every remaining URL was correct: it rendered, had no noindex, did not redirect. Fewer URLs in a sitemap looks like housekeeping, not like a defect. The only way to detect it is to compare the URL set before and after the change.

#Bug 4: >- as an image URL

The integration pulled heroImage out of frontmatter with a regular expression, line by line. With a YAML folded block (heroImage: >-) the value sits on the next line, so the regex captured the literal >-. The sitemap got entries like <image:loc>https://wppoland.com/>-</image:loc>.

On 17 September 2026 production had 31 such entries across six languages. Search Console showed five errors on pl/sitemap-blog.xml, while the XML itself was well formed and every <loc> was fine. The YAML in the files was valid too. The reader was what was broken.

The fix parses frontmatter with a YAML parser, and we exported the metadata extraction function from the integration so a test can reach it without a full build. Before that it sat in the middle of a 900-line file and no test had access to it.

#Bug 5: IndexNow submitted URLs that do not exist

The IndexNow script built URLs from file names in src/content/. On 17 September 2026 it generated 2328 URLs, 0 of which appeared in the sitemap. Blog posts got a /blog/ segment that the real URLs do not have, and pages kept the language suffix from the file name (/de/about.de/). Both variants returned 301. On top of that, the file pattern matched only .md, so 100 service pillars in .mdx were never submitted at all.

The fix fits in one sentence: the source of URLs for IndexNow is sitemap-index.xml. That set is already verified: the pages render and carry no noindex. The file-name heuristic went away, and three classes of bugs went with it.

#A bug without a bug: Google did not fetch the sitemap

This case was not a generator defect, but it belongs to the same family. In early September 2026, after the city pages were unblocked, six locations sitemaps held 4071 URLs. Everything on our side was correct: samples returned 200 and index, follow, and no URL in the sitemap was noindex or a redirect.

Search Console last fetched those sitemaps on 1 September at 10:08, when they had 2037 URLs, and did not refresh them for five days. It read the blog sitemaps daily. Everything added after that time did not exist for Google. A resubmission through the API was fetched the same day.

In the following week the number of city pages with impressions rose from 223 to 343, and their impressions from 1570 to 4382. Those 343 pages are 8.4 percent of 4071, and in seven days they brought 2 clicks. What got unblocked was discovery, not traffic, and those are two different numbers.

#What an entry inspector will not see

BugHow it looked in the sitemapWill an entry check catch itWhat caught it
Build date as lastmodevery entry with one dateyes, as a repeated lastmodnumber of distinct dates in the file
Alternates next to noindexcorrect <loc>, bad xhtml:linkpartly, if it checks alternatesSearch Console noindex report
500 pages outside the sitemapevery present entry correctnocomparing the URL set before and after
>- as an image31 bad image:locyes, if it checks imagessitemap errors in Search Console
IndexNow outside the sitemapsitemap correctno, the bug is outside the sitemapintersecting the submission list with the sitemap
Sitemap not fetchedfile correct and currentnolast read date in Search Console

A plugin like the one Joost used answers the question “is what is in the sitemap correct”. It is a good question, and as he showed, most sites fail it. But three of our six cases were about something else: what is missing from the sitemap, what consumes it, and which version Google sees.

#Checklist for a sitemap generator

An entry check is necessary, not sufficient. These five points check the generator, not just the file:

  1. Count the distinct lastmod values in each file. One value across hundreds of entries means a stamp, not a date. The same applies before substituting a new date field: if almost every file holds the same value in it, the field is unfit for lastmod. No lastmod is better than a false one.
  2. Compare the URL set before and after any generator change. The number of added and removed URLs, with the list. A drop in URLs needs an explanation just as much as a rise.
  3. Every field that carries a URL goes through the same filters. loc, language alternates, x-default, image:loc. A filter on one of them does not protect the others.
  4. A rule copied from another component comes with its condition. A redirect “on 404” is not the same as a redirect always.
  5. The sitemap is the only URL source for everything that passes URLs on: IndexNow, the Indexing API, lists for manual submission. And once a week: the last read date in Search Console and the URL count shown there, next to the number of <loc> elements in the live file.

Two things from Google’s documentation worth keeping at hand: a single sitemap file holds at most 50 000 URLs or 50 MB uncompressed, and <priority> and <changefreq> values are ignored. They do no harm, but they are not worth working on.

If the site runs as headless WordPress, there is an extra question: who generates the sitemap at all, the front end or WordPress. We covered that separately in our piece on the sitemap and canonical URL in headless WordPress.

Next step

Turn the article into an actual implementation

This block strengthens internal linking and gives readers the most relevant next move instead of leaving them at a dead end.

Want this implemented on your site?

If you are planning a Headless WordPress setup, frontend decoupling, or migration to Astro, I can design and build the architecture, API, and frontend.

Related cluster

Explore other WordPress services and knowledge base

Strengthen your business with professional technical support in key areas of the WordPress ecosystem.

Article FAQ

Frequently asked questions

Practical answers to apply the topic in real execution.

SEO-readyGEO-readyAEO-ready4 Q&A
Can lastmod be the build date?#
No. Google uses lastmod only when the value is consistently and verifiably accurate to the page's last change. A build date changes on every deployment, so it teaches Google that the signal means nothing. Leaving lastmod out is better than stating a false one.
Do priority and changefreq in a sitemap matter?#
Not to Google. The Google Search Central documentation says plainly that these values are ignored. They do no harm, but they are not worth your time.
How do I check whether Google sees the current sitemap?#
In Search Console, compare the last read date and the number of discovered URLs with the number of loc elements in the live file. A gap means Google is working from an old version and new pages do not exist for it.
Where should the URL list for IndexNow come from?#
From the finished sitemap, not from file names in the repository. The sitemap is already the set of pages that render and carry no noindex. URLs assembled from file names bypass routes, slugs and redirects.

Need an FAQ tailored to your industry and market? We can build one aligned with your business goals.

Let’s discuss

Related Articles