Skip to content
Developer tools 7 min read By The Crawl Cove team

SEO Crawler CLI: Audit a Site From the Terminal

How to run a technical SEO audit from the command line with crawlcove-cli, an open-source SEO crawler CLI: one npx command, JSON or CSV out, and a CI exit code.

Key takeaways

  • An SEO crawler CLI is for the crawls a desktop app is bad at: scheduled, scripted, and inside a pipeline
  • crawlcove-cli runs with one npx command on Node 18+, nothing installed, and checks status codes, redirect chains, titles, metas, H1s, canonicals, noindex and broken internal links
  • The exit code is the point: choose which checks fail the run and at what count, and a cron job or CI step can act on it
  • It respects robots.txt by default and honours a User-agent: crawlcove-cli group, with an override for hosts you own
  • It is a first pass, not the audit: same-origin only, meta-tag noindex only, and no history, so open the same site in the desktop app for the full picture

Most SEO crawls happen in a window: open the app, paste a URL, wait, read the findings. That is the right shape for an audit you are doing by hand. It is the wrong shape for the crawl you want to happen every night at two, the one you want to run before a release, or the one that should fail a build when someone ships a page with no title. Those want a terminal, an exit code, and no human in the loop. This is a walkthrough of doing that with crawlcove-cli, the open-source SEO crawler CLI behind Crawl Cove's GitHub Action and MCP server.

Run it once, nothing installed

The CLI needs Node 18 or later and nothing else. The quickest way to see what it does is a one-off run straight from GitHub:

npx github:CrawlCove/crawlcove-cli crawl https://example.com

That fetches the CLI, crawls up to 100 same-origin pages starting from the URL you gave it, and prints one JSON object per page to standard output. Progress and the summary go to standard error, so you can pipe the JSON somewhere without the chatter mixed in.

If you want a permanent crawlcove command, install the release tarball globally:

npm install -g https://github.com/CrawlCove/crawlcove-cli/archive/refs/tags/v1.1.2.tar.gz
crawlcove --version

Note

The tarball form is deliberate. A global npm install -g github:CrawlCove/crawlcove-cli looks like it works on npm 10 but links the global package to a temporary clone inside npm's cache, which is cleaned up moments later and leaves a dangling command. The npm package itself is coming; until it is published, use npx for one-off runs and the tarball for a global install.

What one crawl gives you

For every page it reaches, the CLI records the things a first technical pass needs:

  • Status code and any fetch error, so 404s, 500s and timeouts are named.
  • Redirect hops, counted per URL. A single redirect is normal and not flagged; two or more in a row is a chain, and chains are the thing that quietly cost crawl budget.
  • Title, meta description and H1, and whether each is missing.
  • Canonical, exactly as the page declares it.
  • noindex, read from the robots meta tag.
  • Broken internal links, reported as the page that links and the page that fails, so you know which page to edit.

Ask for CSV instead of JSON with --output csv, or write either to a file with --file:

crawlcove crawl https://example.com --max-pages 500 --output csv --file crawl.csv

The JSON field names match the Crawl Cove export spec wherever the CLI checks the same thing the desktop app does, so a script written against one reads the other.

Make it fail when it should

Printing findings is useful. Exiting non-zero is what makes a crawl part of a process. Two flags control that:

  • --fail-on takes a comma-separated list of the checks that count: broken-links, missing-titles, noindex, redirect-chains, or none to report without ever failing. The default is all four.
  • --threshold is how many of those findings it takes. The default is 1.

So a pre-release check that only cares about broken links, and tolerates everything else, looks like this:

crawlcove crawl https://staging.example.com --fail-on broken-links --threshold 1
echo "exit code: $?"

Exit codes are simple on purpose: 0 is clean, 1 means the selected checks reached the threshold, and 2 means a usage error or that robots.txt disallowed the start URL itself, so nothing was fetched. A pipeline should treat 2 as a failure as well, because a crawl that fetched nothing has not proved anything.

A nightly crawl in cron

This is the use a desktop app cannot serve. One line in a crontab crawls the site every night and keeps a dated copy:

0 2 * * * cd /var/crawls && npx -y github:CrawlCove/crawlcove-cli crawl https://example.com --max-pages 1000 --file "$(date +\%F).json" --fail-on none

With --fail-on none the job never fails; it just records. The value is in the diff. Two dated JSON files and a short script tell you which pages gained a redirect hop, lost a title, or started returning 404 between one night and the next, without anyone opening anything. If you want a person alerted, drop --fail-on none, choose the checks that matter, and let the non-zero exit trigger whatever your scheduler does on failure.

robots.txt, and the sites you own

The CLI fetches /robots.txt before its first page request and never fetches a URL that file disallows. It honours a User-agent: crawlcove-cli group if you publish one and falls back to the * group otherwise. A missing robots.txt allows everything; a server error or an unreachable origin allows nothing. Those are the same defaults as the desktop app, and they are the right defaults for a crawler that anyone can point at anyone's site.

The exception is a host you own that blocks all crawlers on purpose, which describes most staging sites. For those, --ignore-robots turns the check off. The flag is for sites you own or are authorised to crawl, and the CLI says so in its own help text.

Tip

If you run this against your own production site on a schedule, publish a User-agent: crawlcove-cli group in robots.txt with a Crawl-delay or a Disallow for the paths you do not need crawled, and use --concurrency 1 on shared hosting. The CLI ignores Crawl-delay in this version, so the concurrency flag is how you pace it. How to write a robots.txt file covers the syntax.

What it does not do, on purpose

A CLI that fits in a pipeline has to stay small, and this one has three limits worth knowing before you rely on it:

  • It reads noindex from the robots meta tag only, not from the X-Robots-Tag response header. A header-level noindex will not be flagged.
  • It follows same-origin links only. There is no way yet to scope a crawl to a URL prefix, and it will not follow a link to a subdomain.
  • It ignores Crawl-delay. Use --concurrency 1 on a site that asks for one.

It also has no memory. Each run is a fresh crawl with no history, no ranking of findings by impact, and no Search Console data next to them. That is the line between a first pass and an audit. When the CLI fails a build and you want to know why it matters, open the same site in the Crawl Cove desktop crawler: it runs the full check set across every page, ranks the findings, keeps history between crawls, and lets you export the whole dataset in the same format the CLI's output follows.

Next steps

  • The CLI is the engine inside the Crawl Cove GitHub Action, which runs the same crawl on every pull request and posts the result as a PR comment. Putting an SEO check in CI walks through that.
  • If you would rather ask questions than read JSON, the MCP server runs the same crawler for Claude, Cursor and Claude Code.
  • Two checks the crawl does not run have CLIs of their own, in the same shape: the hreflang checker validates a localised page's codes and return tags, and the schema validator checks a page's JSON-LD against the properties Google requires. Both take a URL, print JSON with --json, and exit 1 on an error, so they slot into the same cron job or CI step.
  • For what a full technical pass should cover beyond these checks, the technical SEO audit checklist is the map.

Frequently asked questions

What is an SEO crawler CLI?
A crawler you run from the terminal rather than a window. You give it a URL, it fetches the pages it can reach from there, checks each one, and prints the result as text, JSON or CSV. Because it runs without a screen it can be scheduled, scripted and put inside CI, which is where a desktop crawler cannot go.
Do I need to install anything?
Node 18 or later. The one-off form, npx github:CrawlCove/crawlcove-cli crawl <url>, downloads and runs the CLI without installing it. For a permanent crawlcove command, install the release tarball globally; the README explains why the global github: form is avoided on npm 10.
What does it check?
Status codes and fetch errors, redirect chains counted per hop, missing titles, meta descriptions and H1s, the canonical each page declares, noindex from the robots meta tag, and broken internal links reported as the page that links and the page that fails.
Can it fail a deploy?
Yes. That is what --fail-on and --threshold are for. Pick the checks that count as regressions for your site, set how many it takes, and the process exits 1 when they are reached. Exit 2 means nothing could be fetched at all, which a pipeline should treat as a failure too.
Does it respect robots.txt?
By default, yes. It fetches robots.txt before anything else and never requests a URL that file disallows for crawlcove-cli or for *. A staging host behind a blanket Disallow: / needs --ignore-robots, which is only appropriate for a site you own or are authorised to crawl.
How is it different from Crawl Cove itself?
It is the small, scriptable subset. The desktop app crawls the whole site, ranks every finding by impact, keeps history between crawls, and shows Search Console data alongside. The CLI gives you the checks that fit in a text result and an exit code, and its output uses the same field names as the app's export.

Audit your site the easy way

Crawl Cove finds these issues on your machine and tells you exactly what to fix first. See the features or compare the plans.

Download Crawl Cove