Node CLI · v1.0.0 · MIT
A robots.txt tester for the terminal and CI
crawlcove-robots-txt-tester is the command line companion to the robots.txt generator on this site. It fetches a site's robots.txt, lints it for the mistakes that either block everything or do nothing at all, and tests any URL you name against any user-agent token, reporting the line number that decided each verdict. Its headline use is the classic "staging robots.txt shipped to production": run it in a deploy pipeline with --expect-allowed and the deploy fails before search engines notice.
Free · MIT-licensed · README checked September 2026
Install and run
# one-off, nothing installed (Node 18+):
npx github:CrawlCove/crawlcove-robots-txt-tester https://example.com -u /pricing
# global command, from the release tarball:
npm install -g https://github.com/CrawlCove/crawlcove-robots-txt-tester/archive/refs/tags/v1.0.0.tar.gz
robots-txt-tester --version
Options
robots-txt-tester <site> [options]
-u, --url <path> URL or path to test (repeatable; default /)
-a, --agent <token> user-agent token to test as (repeatable; default *, Googlebot, Bingbot)
--expect-allowed fail if any tested URL is blocked for any tested agent
--no-check-sitemaps skip HEAD requests to the Sitemap: URLs
--timeout <ms> per-request timeout (default 10000)
--json JSON output
--fail-on <level> error (default), warning, none
The README in the repo is the authoritative reference; the commands above are copied from it as of September 2026. If they disagree, the README is newer.
What it checks
- blocks-everything: Disallow: / under User-agent: * or Googlebot, the mistake that removes a site from search.
- url-blocked: each tested URL, per agent, with the line that decided it.
- no-sitemap: no Sitemap: line, so crawlers cannot find the sitemap without a Search Console submission.
- Sitemap: URLs are HEAD-requested so a dead sitemap reference is caught too.
- Exit code 0 when clean, 1 on a failing finding, 2 on a usage error, so it slots into any pipeline.
Works with Crawl Cove
robots.txt tester CLI covers the checks it lists above. The desktop app runs the full site audit: every page, every finding ranked by impact, history over time and Search Console data alongside. See the desktop checker or compare the plans.
Frequently asked questions
- What does --expect-allowed do?
- It turns the tester into a guard. If any tested URL is blocked for any tested agent, the command exits 1. Put it in a deploy step with your homepage and a few key paths and a staging robots.txt can never reach production unnoticed.
- Which crawlers can I test as?
- Any user-agent token. By default it tests *, Googlebot and Bingbot; pass -a as many times as you like for others, including AI crawlers.
- Does it write a robots.txt for me?
- No, it tests one. To write or rewrite a robots.txt, use the generator on this site, then run the tester against the published file.
Stop guessing.
Start fixing.
Crawl Cove runs on your machine, connects to your real ranking data, and tells you exactly what to fix first. No per-feature paywalls, no spreadsheets, no guesswork.
28 days risk-free · No card required