Skip to content

Node CLI · v1.0.0 · MIT

An XML sitemap validator for the terminal and CI

crawlcove-sitemap-validator is the command line half of the XML sitemap checker on this site. Give it a sitemap URL, or a site URL and it will read the Sitemap: lines from robots.txt or fall back to /sitemap.xml, and it validates the XML against the sitemaps.org protocol and the limits search engines enforce. It follows a sitemap index into every child, decompresses .gz files, optionally requests the first N listed URLs to flag anything that is not a 200, and exits non-zero on errors so a deploy can catch a broken sitemap before Search Console does.

Free · MIT-licensed · README checked September 2026

Install and run

# one-off, nothing installed (Node 18+):
npx github:CrawlCove/crawlcove-sitemap-validator https://example.com

# global command, from the release tarball:
npm install -g https://github.com/CrawlCove/crawlcove-sitemap-validator/archive/refs/tags/v1.0.0.tar.gz
sitemap-validator --version

Options

sitemap-validator <url> [options]

  <url>                 a sitemap URL, or a site URL (robots.txt Sitemap: lines, else /sitemap.xml)
  --check-urls <n>      also request the first N listed URLs and flag non-200s (max 200)
  --max-sitemaps <n>    child sitemaps of an index to validate (default 50)
  --timeout <ms>        per-request timeout (default 15000)
  --json                JSON output
  --fail-on <level>     error (default), warning, none

The README in the repo is the authoritative reference; the commands above are copied from it as of September 2026. If they disagree, the README is newer.

What it checks

  • not-found, fetch-error and not-xml: the sitemap URL must return real XML with a 200, not an error page, login wall or WAF challenge.
  • cross-host-loc: a URL on a different host than the sitemap, which search engines ignore unless it is cross-submitted via robots.txt.
  • duplicate-loc and empty sitemaps.
  • invalid-lastmod: a date that is not W3C Datetime, and changefreq or priority values outside the protocol.
  • Sitemap index handling: every child validated, gzipped children decompressed, with size and URL counts per file.

Works with Crawl Cove

Sitemap validator CLI covers the checks it lists above. The desktop app runs the full site audit: every page, every finding ranked by impact, history over time and Search Console data alongside. See the desktop checker or compare the plans.

Frequently asked questions

Does it check that the listed URLs actually work?
Optionally. Pass --check-urls with a number and it requests that many of the listed URLs and flags anything that does not return 200, up to 200 URLs. The default run validates the XML only.
What is the most common error it finds?
The README lists each finding with its fix. Two that come up constantly: an unescaped ampersand in a loc, which makes the whole file malformed XML, and lastmod values written as a local date format rather than W3C Datetime, which search engines silently ignore.
How do I find pages that are missing from the sitemap?
That needs a crawl to compare against, which this tool does not do. Crawl Cove crawls the site and compares what it finds with what the sitemap lists, in both directions.

Stop guessing.
Start fixing.

Crawl Cove runs on your machine, connects to your real ranking data, and tells you exactly what to fix first. No per-feature paywalls, no spreadsheets, no guesswork.

28 days risk-free · No card required