On 27 September 2026 Joost de Valk published a piece arguing that almost nobody gets XML sitemaps right. He ran six sites, from Adobe to GOV.UK, through his Sitemap Inspector plugin and found the same things in every one: noindexed URLs, redirects, 404s, canonicals pointing elsewhere, one date across hundreds of entries. His thesis: “A sitemap should list the URLs a site wants indexed, and nothing else.”
We agree with the thesis and with the list. We used it on our own site. Between August and October 2026 we found five bugs in the wppoland.com sitemap generator, each one measured in production. Joost’s list would have caught three of them. No inspector that reads sitemap entries will catch the other two, because they were about what the sitemap did not contain. This piece covers both kinds.
Technical context: wppoland.com is an Astro site in six languages, with about 6700 URLs across its sitemaps. The sitemaps come from a custom integration, astro-sitemaps.mjs, not a plugin. That matters for the conclusions: every bug below lived in the code that reads data and builds the list, not in the list itself.
Bug 1: build date as lastmod
On 24 August 2026, five different sitemaps had exactly one lastmod value for every entry, the same one, today’s. The integration computed new Date() once and stamped it on 3079 URLs. Every deployment told Google that everything had changed.
Google states the condition directly: it uses <lastmod> if the value is “consistently and verifiably” accurate to the page’s last change (Google Search Central, the page on building sitemaps, updated 8 July 2026). A date that jumps with every build fails that test, so Google stops trusting it for the whole domain, including where it would have been accurate. In samples from that period, the last crawl of some pages went back anywhere from three weeks to three months.
The fix substituted updatedDate from frontmatter, falling back to pubDate. Effect on the artifact: the Polish blog sitemap went from 1 to 98 distinct dates.
This fix invites a second-order mistake. It was tempting to use the lastVerified field for city pages, since every one of the 4733 files has it. Except 4697 of the 4733 files hold the same value in it. That is a bulk stamp, not a signal, and it would have moved the same defect from “today” to a different constant date. Before any date field goes into lastmod, count how many distinct values it has.
The August fix covered the blog, the portfolio and pages from the content collection. It did not cover hand-written .astro routes. Joost’s article prompted us to check again: on 29 September, in /pl/sitemap-pages.xml, 62 of 113 URLs still had the build date as lastmod, because no content date exists for them and the code then fell back to today. The fix (commit af97c7a887) takes the date of the source file’s last commit for such routes, from a single git log call per build, and when there is none, omits lastmod. The shared template [lang]/[slug].astro deliberately does not count as a source, because its date says nothing about any particular page. The same file after the fix: 87 of 113 URLs have lastmod, 42 distinct dates, the most common shared by 7 URLs, and 26 URLs have no date at all. In the Polish city pages sitemap, 684 of 692 entries carried the build date before the fix; after it, 8 entries have a real date: the city template has no date that could honestly be used, so the rest declare none. The fix has been in production since 29 September 2026, and the live file shows the same figures.
Bug 2: the noindex filter guarded one field
After 723 city pages were demoted to noindex, Search Console kept reporting them as discovered. The pages themselves had a correct noindex. They were getting back to Google another way: sitemap-locations.xml listed them as language alternates (xhtml:link, including x-default) of pages that stayed in the index.
The integration’s noindex filter checked only <loc>. The alternates come from the page head, which lists every language version, so the filter never saw them. The fix keeps only the alternates whose URL is itself a <loc> in the sitemap. After it, the build produced 6407 URLs and 0 bad alternates.
The lesson goes beyond sitemaps: a filter on one field of an entry does not cover the other fields of that entry that carry a URL. A sitemap has several such places: loc, alternates, x-default, image:loc.
Bug 3: 500 live pages left out of the sitemap
On 21 September 2026 the integration started skipping every URL matching the prefixes that the middleware in functions/_middleware.ts redirects. But the middleware redirects them only when the resource returns 404, and every sitemap candidate is a built page, so it never returns 404. The filter copied the rule without its condition.
Result: 500 live pages with index, follow dropped out of the sitemap, including the target of one of the redirects (/pl/audyt-bezpieczenstwa-wordpress/ matched the prefix audyt-bezpieczenstwa-). The fix landed on 26 September, and the build with the sitemap gave +500 URLs and 0 removed.
No entry check will catch this bug. Every remaining URL was correct: it rendered, had no noindex, did not redirect. Fewer URLs in a sitemap looks like housekeeping, not like a defect. The only way to detect it is to compare the URL set before and after the change.
Bug 4: >- as an image URL
The integration pulled heroImage out of frontmatter with a regular expression, line by line. With a YAML folded block (heroImage: >-) the value sits on the next line, so the regex captured the literal >-. The sitemap got entries like <image:loc>https://wppoland.com/>-</image:loc>.
On 17 September 2026 production had 31 such entries across six languages. Search Console showed five errors on pl/sitemap-blog.xml, while the XML itself was well formed and every <loc> was fine. The YAML in the files was valid too. The reader was what was broken.
The fix parses frontmatter with a YAML parser, and we exported the metadata extraction function from the integration so a test can reach it without a full build. Before that it sat in the middle of a 900-line file and no test had access to it.
Bug 5: IndexNow submitted URLs that do not exist
The IndexNow script built URLs from file names in src/content/. On 17 September 2026 it generated 2328 URLs, 0 of which appeared in the sitemap. Blog posts got a /blog/ segment that the real URLs do not have, and pages kept the language suffix from the file name (/de/about.de/). Both variants returned 301. On top of that, the file pattern matched only .md, so 100 service pillars in .mdx were never submitted at all.
The fix fits in one sentence: the source of URLs for IndexNow is sitemap-index.xml. That set is already verified: the pages render and carry no noindex. The file-name heuristic went away, and three classes of bugs went with it.
A bug without a bug: Google did not fetch the sitemap
This case was not a generator defect, but it belongs to the same family. In early September 2026, after the city pages were unblocked, six locations sitemaps held 4071 URLs. Everything on our side was correct: samples returned 200 and index, follow, and no URL in the sitemap was noindex or a redirect.
Search Console last fetched those sitemaps on 1 September at 10:08, when they had 2037 URLs, and did not refresh them for five days. It read the blog sitemaps daily. Everything added after that time did not exist for Google. A resubmission through the API was fetched the same day.
In the following week the number of city pages with impressions rose from 223 to 343, and their impressions from 1570 to 4382. Those 343 pages are 8.4 percent of 4071, and in seven days they brought 2 clicks. What got unblocked was discovery, not traffic, and those are two different numbers.
What an entry inspector will not see
| Bug | How it looked in the sitemap | Will an entry check catch it | What caught it |
|---|---|---|---|
| Build date as lastmod | every entry with one date | yes, as a repeated lastmod | number of distinct dates in the file |
| Alternates next to noindex | correct <loc>, bad xhtml:link | partly, if it checks alternates | Search Console noindex report |
| 500 pages outside the sitemap | every present entry correct | no | comparing the URL set before and after |
>- as an image | 31 bad image:loc | yes, if it checks images | sitemap errors in Search Console |
| IndexNow outside the sitemap | sitemap correct | no, the bug is outside the sitemap | intersecting the submission list with the sitemap |
| Sitemap not fetched | file correct and current | no | last read date in Search Console |
A plugin like the one Joost used answers the question “is what is in the sitemap correct”. It is a good question, and as he showed, most sites fail it. But three of our six cases were about something else: what is missing from the sitemap, what consumes it, and which version Google sees.
Checklist for a sitemap generator
An entry check is necessary, not sufficient. These five points check the generator, not just the file:
- Count the distinct
lastmodvalues in each file. One value across hundreds of entries means a stamp, not a date. The same applies before substituting a new date field: if almost every file holds the same value in it, the field is unfit forlastmod. Nolastmodis better than a false one. - Compare the URL set before and after any generator change. The number of added and removed URLs, with the list. A drop in URLs needs an explanation just as much as a rise.
- Every field that carries a URL goes through the same filters.
loc, language alternates,x-default,image:loc. A filter on one of them does not protect the others. - A rule copied from another component comes with its condition. A redirect “on 404” is not the same as a redirect always.
- The sitemap is the only URL source for everything that passes URLs on: IndexNow, the Indexing API, lists for manual submission. And once a week: the last read date in Search Console and the URL count shown there, next to the number of
<loc>elements in the live file.
Two things from Google’s documentation worth keeping at hand: a single sitemap file holds at most 50 000 URLs or 50 MB uncompressed, and <priority> and <changefreq> values are ignored. They do no harm, but they are not worth working on.
If the site runs as headless WordPress, there is an extra question: who generates the sitemap at all, the front end or WordPress. We covered that separately in our piece on the sitemap and canonical URL in headless WordPress.







