Skip to content

Node CLI · v1.1.2 · MIT

A command line site crawler for scripts and CI

crawlcove-cli crawls a site from the terminal and reports the technical checks a first SEO pass needs: status codes, redirect chains, titles, meta descriptions, H1s, canonicals, noindex and broken internal links. It is the crawler behind the Crawl Cove GitHub Action and the MCP server, and its per-page output uses the same field names as the export spec, so the three fit together.

Free · MIT-licensed · README checked September 2026

Install and run

# one-off, nothing installed (Node 18+):
npx github:CrawlCove/crawlcove-cli crawl https://example.com

# global command, from a release tarball:
npm install -g https://github.com/CrawlCove/crawlcove-cli/archive/refs/tags/v1.1.2.tar.gz
crawlcove --version

Options

crawlcove crawl <url> [options]

  -o, --output <format>   json or csv (default: json)
  --file <path>           write output to a file instead of stdout
  --max-pages <n>         stop after this many pages (default: 100)
  --concurrency <n>       simultaneous requests (default: 4)
  --timeout <ms>          per-request timeout (default: 15000)
  --ignore-robots         crawl URLs robots.txt disallows (sites you own only)
  --fail-on <checks>      broken-links, missing-titles, noindex, redirect-chains, or none
  --threshold <n>         exit non-zero once the selected checks total this many (default: 1)

The README in the repo is the authoritative reference; the commands above are copied from it as of September 2026. If they disagree, the README is newer.

What it checks

  • Status codes and fetch errors for every same-origin page it reaches.
  • Redirect chains, counted per hop, so a two-hop chain is flagged and a single redirect is not.
  • Missing titles, meta descriptions and H1s, plus the canonical each page declares.
  • noindex, read from the robots meta tag.
  • Broken internal links, reported as the page that links and the page that fails.
  • Exit code 0 when clean, 1 when the checks you chose reach the threshold, 2 when nothing could be fetched.

Limits worth knowing

  • Reads noindex from the robots meta tag only, not the X-Robots-Tag header.
  • Follows same-origin links only; there is no URL-prefix scoping yet.
  • Ignores Crawl-delay, so use --concurrency 1 on a site that asks for one.

robots.txt and the crawlcove-cli user agent

The CLI fetches robots.txt before anything else and never requests a URL it disallows. It honours a User-agent: crawlcove-cli group if you publish one, otherwise the * group. A missing robots.txt allows everything; a server error or unreachable origin allows nothing, the same defaults as the desktop app.

Seen it in your logs? Someone ran this crawler against your site from their own machine. It reads HTML and reports problems; it submits nothing. The desktop app's crawler page covers the same ground for the Crawl Cove desktop user agent.

Works with Crawl Cove

Crawl Cove CLI covers the checks it lists above. The desktop app runs the full site audit: every page, every finding ranked by impact, history over time and Search Console data alongside. See the desktop checker or compare the plans.

Frequently asked questions

Is this the same crawler as the Crawl Cove desktop app?
No. It is a separate, smaller crawler that shares the desktop app's checks and output shape for the things it covers. The desktop app runs the full audit: every page, every finding ranked by impact, history over time and Search Console data alongside. The CLI is for scripted, scheduled or CI crawls where you want a text result and an exit code.
Does it respect robots.txt?
Yes, by default. It fetches robots.txt first and skips anything disallowed for crawlcove-cli or *, listing skipped URLs in the output. The --ignore-robots flag exists for hosts you own, such as a staging site behind a blanket Disallow.
Can I install it with npm install -g crawlcove?
Not yet. The npm package is coming; until then use npx straight from GitHub for a one-off run, or install the release tarball globally. The README explains why a global github: install is avoided on npm 10.
I saw crawlcove-cli in my server logs. What is it?
Someone ran this open-source crawler against your site from their own machine, most likely the owner, a developer or an agency working on it. It reads HTML and reports problems; it submits nothing. Add a User-agent: crawlcove-cli group to your robots.txt to slow it down or block it.

Stop guessing.
Start fixing.

Crawl Cove runs on your machine, connects to your real ranking data, and tells you exactly what to fix first. No per-feature paywalls, no spreadsheets, no guesswork.

28 days risk-free · No card required