Key takeaways
- An SEO GitHub Action moves the crawl from after the release to before the merge, where a broken link or an accidental noindex is a failed check instead of a lost week of traffic
- crawlcove-action is three lines in a workflow: uses, with, url, and it crawls up to 100 pages and posts one PR comment it keeps updating
- Choose which checks fail the build with fail-on: previews are often noindex on purpose, so drop that check there rather than everywhere
- Hosted previews arrive via a deployment_status event with no secrets, a Next.js build can be served inside the job, and a scheduled WordPress staging crawl needs no PR at all
- The action catches regressions; it does not replace the audit, and its report follows the same field names as the desktop app's export
The expensive SEO mistakes are rarely clever. A template change drops the title tag from every product page. A redirect rule gets a second hop. Someone ships the staging noindex to production. None of these are hard to find with a crawler; the problem is that the crawl happens after the release, when the damage is already in the index. crawlcove-action moves that crawl to the pull request, where a regression is a red check instead of a lost week of traffic. This is how to set it up, and how to choose what should fail.
The three-line version
The action wraps the Crawl Cove CLI. The smallest useful workflow step is:
- uses: CrawlCove/crawlcove-action@v1
with:
url: https://preview-123.example.com/
That crawls up to 100 same-origin pages from the URL, fails the check on the first broken internal link, missing title, noindex page or two-hop redirect chain, and posts one comment on the pull request that it updates on every run rather than duplicating. The comment looks like this:
CrawlCove SEO crawl: failed Crawled 42 pages from
https://preview-123.example.com/. The gated checks total 3, threshold 1.
Check Count Fails the build Broken links 2 yes Missing titles 1 yes Noindex pages 0 yes Redirect chains (2+ hops) 0 yes
The broken links and missing titles expand below the table, as the page that links and the page that fails, so the fix is one click away rather than a search.
The comment needs permissions: pull-requests: write on the job. Everything else works with the default token.
Getting a URL to crawl
The action is only as useful as the URL you give it, and there are three common shapes.
A hosted preview (Vercel, Netlify, Cloudflare Pages and similar). The host's GitHub integration sends a deployment_status event when a preview is up, and the event carries the URL. No secrets, no waiting loop:
name: SEO crawl (preview)
on:
deployment_status:
permissions:
contents: read
pull-requests: write
jobs:
crawl:
if: github.event.deployment_status.state == 'success'
runs-on: ubuntu-latest
steps:
- uses: CrawlCove/crawlcove-action@v1
with:
url: ${{ github.event.deployment_status.environment_url }}
max-pages: 200
fail-on: broken-links,missing-titles,redirect-chains
Note the fail-on list: noindex is left out, because preview deployments are usually noindex on purpose. More on that below.
A site built inside the job. For a Next.js site, or anything that can build and serve on localhost, there is no external host to wait for. Build, serve, wait for the port, crawl:
on:
pull_request:
jobs:
crawl:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
cache: npm
- run: npm ci
- run: npm run build
- name: Serve the build
run: |
(npx next start -p 3000 > /dev/null 2>&1 &)
for i in $(seq 1 30); do curl -fs http://127.0.0.1:3000/ > /dev/null && break; sleep 1; done
- uses: CrawlCove/crawlcove-action@v1
with:
url: http://127.0.0.1:3000/
max-pages: 300
ignore-robots: 'true'
ignore-robots is set because a local build's robots.txt usually blocks everything, and a crawler that respects it would fetch nothing. Which brings up the third shape.
A staging site on a schedule. A WordPress staging install does not have pull requests, but it has a URL and a cron schedule. The repo's wordpress-staging.yml example runs the same crawl on a timer with comment off, so a nightly regression shows up as a failed workflow run rather than a PR comment.
Choosing what fails
Four checks are gated by default, and the right list depends on what the URL is.
- broken-links counts every link to a 4xx, 5xx or unreachable page, plus each such page itself. This is the one to keep on everywhere. A broken internal link is a regression on any environment.
- missing-titles is a page with no
<title>at all. Also keep on. It is almost always a template bug, and template bugs multiply. - noindex is a page carrying a robots
noindexmeta tag. On production this is the check that catches the shipped staging tag. On a preview it is noise, because the preview is meant to be noindex. Setfail-onper environment rather than switching the check off globally. - redirect-chains counts URLs that went through two or more redirects. A single redirect is normal and is not flagged. Chains cost crawl budget and leak link equity, and they tend to appear exactly when a URL structure changes, which is exactly when a PR is open.
threshold sets how many findings across the selected checks it takes to fail, and defaults to 1. none for fail-on reports without ever failing, which is a reasonable first week while you see what the crawl finds on your site.
Reading the result from later steps
Every count is an output, so a later step can act on it without parsing the comment:
- uses: CrawlCove/crawlcove-action@v1
id: seo
with:
url: ${{ github.event.deployment_status.environment_url }}
- run: echo "Broken links: ${{ steps.seo.outputs.broken-links }}"
- uses: actions/upload-artifact@v4
with:
name: seo-crawl
path: ${{ steps.seo.outputs.report-path }}
report-path points at the full JSON report, one object per page, with field names that follow the Crawl Cove export spec. Anything that reads a Crawl Cove export reads this file.
exit-code is worth knowing too: 0 is clean, 1 is the gated checks reaching the threshold, and 2 means the crawl could not run, most often because robots.txt disallowed the start URL. The action deliberately reports 2 as a failure with a "could not run" summary rather than as a pass, because a crawl that fetched nothing has proved nothing.
Note
cli-version pins which git ref of crawlcove-cli the action installs, and defaults to the CLI release the action was tested with. Leave it alone unless you are testing a newer CLI on purpose.
What a CI crawl is not
A check on a pull request catches regressions on the pages the preview happens to have. It does not tell you which of the findings across a whole site matter most, it keeps no history between runs, and it knows nothing about how the pages perform in search. That is the audit, and it is a different job. When the check goes red and you want the whole picture, open the same site in the Crawl Cove desktop crawler: every page, every finding ranked by impact, history over time and Search Console data alongside. The action is the tripwire; the app is the investigation.
Related
- Run an SEO audit from the terminal covers the CLI the action runs, including cron crawls and the robots.txt rules.
- Redirect chains and loops explains why the two-hop rule is the right cut-off.
- Site migration SEO checklist is the manual version of what this action automates on the day a URL structure changes.
Frequently asked questions
- What is an SEO GitHub Action?
- A step in a GitHub Actions workflow that crawls a URL and fails the check when it finds SEO regressions. crawlcove-action crawls the preview or staging URL of a pull request and fails on broken internal links, missing titles, noindex pages or two-hop redirect chains, posting the detail as a PR comment.
- What does it need from my repository?
- A URL the runner can reach and, if you want the PR comment, pull-requests write permission on the job. Hosted previews supply the URL through the deployment_status event from your host's GitHub integration, so no secrets are involved.
- Will it fail on a preview that is noindex on purpose?
- With the default fail-on list, yes, because noindex is one of the four gated checks. Preview deployments are often noindex deliberately, so set fail-on to broken-links,missing-titles,redirect-chains for those and keep the noindex check for staging or production crawls.
- What happens if robots.txt blocks the crawler?
- The run reports exit code 2 and a could-not-run summary rather than a false pass. For a host you own that blocks all crawlers, such as a local build or a locked-down staging site, set ignore-robots to true.
- Where does the full report go?
- The action writes a JSON report with one object per page and exposes its path as the report-path output. Upload it with actions/upload-artifact to keep it beyond the run. Its page fields follow the Crawl Cove export spec.
- Is it on the GitHub Marketplace?
- Not yet. It works from its repository today with uses CrawlCove/crawlcove-action@v1; a Marketplace listing is a separate submission that has not been made.