Skip to content
Developer tools 6 min read By The Crawl Cove team

SEO GitHub Action: Fail the Build on Broken Links

How to add an SEO check to every pull request with crawlcove-action, an open-source SEO GitHub Action that crawls your preview URL and fails on broken links.

Key takeaways

  • An SEO GitHub Action moves the crawl from after the release to before the merge, where a broken link or an accidental noindex is a failed check instead of a lost week of traffic
  • crawlcove-action is three lines in a workflow: uses, with, url, and it crawls up to 100 pages and posts one PR comment it keeps updating
  • Choose which checks fail the build with fail-on: previews are often noindex on purpose, so drop that check there rather than everywhere
  • Hosted previews arrive via a deployment_status event with no secrets, a Next.js build can be served inside the job, and a scheduled WordPress staging crawl needs no PR at all
  • The action catches regressions; it does not replace the audit, and its report follows the same field names as the desktop app's export

The expensive SEO mistakes are rarely clever. A template change drops the title tag from every product page. A redirect rule gets a second hop. Someone ships the staging noindex to production. None of these are hard to find with a crawler; the problem is that the crawl happens after the release, when the damage is already in the index. crawlcove-action moves that crawl to the pull request, where a regression is a red check instead of a lost week of traffic. This is how to set it up, and how to choose what should fail.

The three-line version

The action wraps the Crawl Cove CLI. The smallest useful workflow step is:

- uses: CrawlCove/crawlcove-action@v1
  with:
    url: https://preview-123.example.com/

That crawls up to 100 same-origin pages from the URL, fails the check on the first broken internal link, missing title, noindex page or two-hop redirect chain, and posts one comment on the pull request that it updates on every run rather than duplicating. The comment looks like this:

CrawlCove SEO crawl: failed Crawled 42 pages from https://preview-123.example.com/. The gated checks total 3, threshold 1.

Check Count Fails the build
Broken links 2 yes
Missing titles 1 yes
Noindex pages 0 yes
Redirect chains (2+ hops) 0 yes

The broken links and missing titles expand below the table, as the page that links and the page that fails, so the fix is one click away rather than a search.

The comment needs permissions: pull-requests: write on the job. Everything else works with the default token.

Getting a URL to crawl

The action is only as useful as the URL you give it, and there are three common shapes.

A hosted preview (Vercel, Netlify, Cloudflare Pages and similar). The host's GitHub integration sends a deployment_status event when a preview is up, and the event carries the URL. No secrets, no waiting loop:

name: SEO crawl (preview)

on:
  deployment_status:

permissions:
  contents: read
  pull-requests: write

jobs:
  crawl:
    if: github.event.deployment_status.state == 'success'
    runs-on: ubuntu-latest
    steps:
      - uses: CrawlCove/crawlcove-action@v1
        with:
          url: ${{ github.event.deployment_status.environment_url }}
          max-pages: 200
          fail-on: broken-links,missing-titles,redirect-chains

Note the fail-on list: noindex is left out, because preview deployments are usually noindex on purpose. More on that below.

A site built inside the job. For a Next.js site, or anything that can build and serve on localhost, there is no external host to wait for. Build, serve, wait for the port, crawl:

on:
  pull_request:

jobs:
  crawl:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 20
          cache: npm
      - run: npm ci
      - run: npm run build
      - name: Serve the build
        run: |
          (npx next start -p 3000 > /dev/null 2>&1 &)
          for i in $(seq 1 30); do curl -fs http://127.0.0.1:3000/ > /dev/null && break; sleep 1; done
      - uses: CrawlCove/crawlcove-action@v1
        with:
          url: http://127.0.0.1:3000/
          max-pages: 300
          ignore-robots: 'true'

ignore-robots is set because a local build's robots.txt usually blocks everything, and a crawler that respects it would fetch nothing. Which brings up the third shape.

A staging site on a schedule. A WordPress staging install does not have pull requests, but it has a URL and a cron schedule. The repo's wordpress-staging.yml example runs the same crawl on a timer with comment off, so a nightly regression shows up as a failed workflow run rather than a PR comment.

Choosing what fails

Four checks are gated by default, and the right list depends on what the URL is.

  • broken-links counts every link to a 4xx, 5xx or unreachable page, plus each such page itself. This is the one to keep on everywhere. A broken internal link is a regression on any environment.
  • missing-titles is a page with no <title> at all. Also keep on. It is almost always a template bug, and template bugs multiply.
  • noindex is a page carrying a robots noindex meta tag. On production this is the check that catches the shipped staging tag. On a preview it is noise, because the preview is meant to be noindex. Set fail-on per environment rather than switching the check off globally.
  • redirect-chains counts URLs that went through two or more redirects. A single redirect is normal and is not flagged. Chains cost crawl budget and leak link equity, and they tend to appear exactly when a URL structure changes, which is exactly when a PR is open.

threshold sets how many findings across the selected checks it takes to fail, and defaults to 1. none for fail-on reports without ever failing, which is a reasonable first week while you see what the crawl finds on your site.

Reading the result from later steps

Every count is an output, so a later step can act on it without parsing the comment:

- uses: CrawlCove/crawlcove-action@v1
  id: seo
  with:
    url: ${{ github.event.deployment_status.environment_url }}
- run: echo "Broken links: ${{ steps.seo.outputs.broken-links }}"
- uses: actions/upload-artifact@v4
  with:
    name: seo-crawl
    path: ${{ steps.seo.outputs.report-path }}

report-path points at the full JSON report, one object per page, with field names that follow the Crawl Cove export spec. Anything that reads a Crawl Cove export reads this file.

exit-code is worth knowing too: 0 is clean, 1 is the gated checks reaching the threshold, and 2 means the crawl could not run, most often because robots.txt disallowed the start URL. The action deliberately reports 2 as a failure with a "could not run" summary rather than as a pass, because a crawl that fetched nothing has proved nothing.

Note

cli-version pins which git ref of crawlcove-cli the action installs, and defaults to the CLI release the action was tested with. Leave it alone unless you are testing a newer CLI on purpose.

What a CI crawl is not

A check on a pull request catches regressions on the pages the preview happens to have. It does not tell you which of the findings across a whole site matter most, it keeps no history between runs, and it knows nothing about how the pages perform in search. That is the audit, and it is a different job. When the check goes red and you want the whole picture, open the same site in the Crawl Cove desktop crawler: every page, every finding ranked by impact, history over time and Search Console data alongside. The action is the tripwire; the app is the investigation.

Frequently asked questions

What is an SEO GitHub Action?
A step in a GitHub Actions workflow that crawls a URL and fails the check when it finds SEO regressions. crawlcove-action crawls the preview or staging URL of a pull request and fails on broken internal links, missing titles, noindex pages or two-hop redirect chains, posting the detail as a PR comment.
What does it need from my repository?
A URL the runner can reach and, if you want the PR comment, pull-requests write permission on the job. Hosted previews supply the URL through the deployment_status event from your host's GitHub integration, so no secrets are involved.
Will it fail on a preview that is noindex on purpose?
With the default fail-on list, yes, because noindex is one of the four gated checks. Preview deployments are often noindex deliberately, so set fail-on to broken-links,missing-titles,redirect-chains for those and keep the noindex check for staging or production crawls.
What happens if robots.txt blocks the crawler?
The run reports exit code 2 and a could-not-run summary rather than a false pass. For a host you own that blocks all crawlers, such as a local build or a locked-down staging site, set ignore-robots to true.
Where does the full report go?
The action writes a JSON report with one object per page and exposes its path as the report-path output. Upload it with actions/upload-artifact to keep it beyond the run. Its page fields follow the Crawl Cove export spec.
Is it on the GitHub Marketplace?
Not yet. It works from its repository today with uses CrawlCove/crawlcove-action@v1; a Marketplace listing is a separate submission that has not been made.

Audit your site the easy way

Crawl Cove finds these issues on your machine and tells you exactly what to fix first. See the features or compare the plans.

Download Crawl Cove