Skip to content
Reference 11 min read

The complete audit-check reference (49 checks)

Every one of Crawl Cove's 49 deterministic checks, grouped by category. Each entry lists id, default severity, what it detects, and its tunable thresholds.

When a client asks "what exactly do you test?", this is the answer, and the detail behind the On-Page SEO Checker landing page. Crawl Cove ships 49 deterministic checks across nine categories: HTTP, meta tags, content, links, security, images, an auto-detected WordPress pack, AI-readiness checks, and performance (Core Web Vitals).

Every check has a stable id, a semantic version, a default severity, and many expose tunable thresholds you can adjust per site in the Check Registry. Findings are keyed by (check id, URL), which is what lets the app compare runs and show honest deltas over time.

Note

Findings are deterministic, never guessed. Every row below is produced by a coded check function, not by an AI. The optional LLM layer only explains a finding in plain English; it can never invent, remove, or rescore one. When you put a number in front of a client, it came from a rule you can point to.

Why the version matters: when you improve a check, its semantic version changes, so run-to-run deltas stay honest rather than silently shifting because the rule moved under you. And because findings are keyed by (check id, URL), the same issue on the same page is tracked as one continuous thread across audits. That thread is the foundation of proving a fix worked.

Note

Accurate as of Crawl Cove 0.1.8 (September 2026). This reference is checked against the check registry in the shipping release. Crawl Cove is actively developed, so future updates are likely to add new checks and categories, and existing checks are sharpened between releases. Treat the count above as today's snapshot, not a permanent ceiling. Your installed version's Check Registry is always the authoritative, up-to-date list.

Note

What "full audits only" means. A few checks reach their conclusion from something being absent, so they only run when the crawl saw everything it found. Crawl Cove treats a crawl as partial if you scoped it with include or exclude patterns, if Max pages (ships at 500) or Max depth (ships at 5) turned an in-scope URL away, or if fewer than 90% of the URLs it discovered were fetched. On a partial crawl broken-internal-link is withheld entirely, and canonical drops only its "target not found" case while its missing and malformed cases carry on: a link or canonical pointing at a page the crawl never reached cannot be told apart from one pointing nowhere. orphan-page and weakly-linked-page soften their severity instead of going quiet, and the two sitemap checks have always needed a full audit. Every audit states on its own screens how much of the site it covered, so a withheld check is never a silent one.

HTTP

Check id Default severity Detects Tunable
status-code Critical Pages that failed to fetch or returned 4xx/5xx (redirects are out of scope). A 429 is the one exception: it is reported at Low, saying the page was not checked, because the server was throttling the crawler rather than serving a broken page. none
redirect-chain Low Pages reached only after one or more redirects, showing the full hop chain. min hops (1)

Meta tags

Check id Default severity Detects Tunable
title High Missing/empty or wrong-length <title>. min 30 / max 60 chars
duplicate-title High Indexable pages sharing a title with other indexable pages. none
meta-description Medium Missing/empty or wrong-length meta description. min 70 / max 160 chars
duplicate-meta-description Medium Indexable pages sharing a meta description. none
canonical Medium Missing or non-absolute canonical on indexable pages. Its third case, a same-origin canonical pointing at a URL the crawl never saw, is full audits only. none
canonical-to-noindex High Canonical resolves to a page that is itself noindex or non-200. none
canonical-to-redirect Medium Canonical points at a redirecting URL, not the redirect's final destination. none
insecure-canonical Medium HTTPS page canonicalised to its http:// version. none
social-meta Low Indexable pages missing Open Graph / Twitter Card tags used for social and AI share previews. none
html-lang Low Missing or invalid root <html lang> attribute. none
hreflang Medium Invalid hreflang codes, missing self-reference, missing return links, or alternates pointing at non-indexable URLs. none
robots-noindex High Surfaces pages excluded from search by a robots directive (noindex, or its synonym none), from either the robots meta tag or the X-Robots-Tag response header, so the exclusion can be confirmed as intentional. none
structured-data-missing Low Substantive indexable pages with no JSON-LD structured data. min 100 words
url-consistency Low Inconsistent casing, trailing slashes, or redundant query parameters that risk duplicate-content dilution. none
sitemap-non-indexable-url Medium Sitemap URLs that are not canonical, indexable 200 pages (uncrawled, redirecting, non-200, or noindex). Full audits only. none
page-missing-from-sitemap Low Indexable, self-canonical crawled pages not listed in the XML sitemap. Full audits only. none

Content

Check id Default severity Detects Tunable
heading-h1 High Not exactly one H1 on the page. expected H1 count (1)
heading-order Low A skipped heading level (e.g. H2 → H4). Readability and accessibility, not ranking. max heading level jump (1)
thin-content Low Indexable pages below the word-count floor; near-empty pages flagged at higher severity. min 200 words, very-thin at 60
content-under-structured Low Substantial pages (≥ 600 words) with fewer than 2 H2-H6 subheadings. min words (600), min subheadings (2)
near-duplicate-content Medium Pages whose main body content is near-identical to other indexable pages (SimHash distance). max fingerprint distance (10 bits)
content-depends-on-js Medium Critical content or structured data that only appears after JavaScript runs (raw-vs-rendered diff). none
schema-validity Medium Malformed JSON-LD, or structured data missing required Schema.org properties. none

Links

Check id Default severity Detects Tunable
broken-internal-link High Internal links to same-site URLs not found in the crawl (typo, deleted page, wrong path). Full audits only. none
orphan-page Medium* Indexable non-seed pages with zero internal inbound links. none
internal-link-to-redirect Medium Internal links pointing at a redirect, not the final URL. none
excessive-crawl-depth Low Indexable pages too many clicks from the seed. max depth (4)
nofollow-internal-link Low Internal links marked rel="nofollow", which block internal PageRank flow. none
dead-end-page Low Indexable pages with no outbound internal links. none
weakly-linked-page Low* Pages with some but fewer than the recommended inbound links (complements orphan-page). min inbound (2)
internal-link-opportunity Info Suggests missing contextual internal links to topically related pages (an opportunity, not a defect). shared terms (2), relatedness (0.18), max/page (5), boilerplate cutoff (0.5)
non-descriptive-anchor-text Low Generic anchors ("click here", "read more"). none
excessive-internal-links Low Pages whose internal link count exceeds the recommended maximum (sprawl). max links (300)

*Severity softens on partial/scoped crawls.

Note

The two checks marked with an asterisk, orphan-page and weakly-linked-page, automatically soften their severity on partial or scoped crawls. Inbound links may well exist on pages you chose not to crawl, so the app refuses to cry "orphan" with full confidence when it can't see the whole link graph. Honesty over false alarms.

Security

Check id Default severity Detects Tunable
mixed-content High HTTPS pages referencing insecure http:// resources in image srcs, link hrefs, or JSON-LD. none

Images

Check id Default severity Detects Tunable
image-alt Medium Pages with images whose alt attribute is absent entirely, or present but only whitespace. A deliberate alt="" is correct markup for a decorative image and is not flagged. min images missing (1)

WordPress (only on auto-detected WordPress sites)

These checks run only when the crawler auto-detects WordPress. A non-WordPress site gets zero WP probes and zero WP findings. See How the crawler behaves.

Check id Default severity Detects Tunable
wp-generator-version-exposed Low WordPress version leaked via the generator meta tag. none
wp-xmlrpc-reachable Medium /xmlrpc.php publicly reachable (DDoS/brute-force vector). none
wp-default-permalinks Low Default ?p=N permalinks instead of descriptive URLs. none
wp-uploads-listing-open Medium /wp-content/uploads/ returns an open directory index. none
wp-missing-image-dimensions Low Images without width/height attributes (causes layout shift). none

AI-readiness

Check id Default severity Detects Tunable
ai-bot-access Medium AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) blocked in robots.txt. none
ai-llms-txt Low No /llms.txt content-map file present (advisory). none
ai-extractability Low Pages structurally hard for AI engines to quote (names the missing signals). none
ai-author-authority Low Article structured data missing an author or author identity link (the E-E-A-T signal generative engines use). none
ai-content-freshness Low Article structured data missing a dateModified freshness signal. none
ai-answer-structure Info Substantial content pages with no question/answer structure for generative engines to extract. min 400 words
ai-entity-markup Low No Organization/Person entity markup with sameAs links for entity disambiguation (advisory). none

Performance

Check id Default severity Detects Tunable
performance-core-web-vitals Medium Pages whose Core Web Vitals are in Google's poor range, from CrUX field data (preferred) or Lighthouse lab data (fallback). LCP (4000 ms), INP (500 ms), CLS (0.25)

Tuning these to a client

Every threshold above is a per-site default, not a hard-coded law. An e-commerce site with faceted URLs needs different link rules than a blog. In the Check Registry you can enable or disable any check per site, override its severity, and adjust its thresholds, all respected by the audit pipeline on the next crawl. See Tuning checks with the Check Registry.

Next

Put this guide into practice

Crawl Cove runs these audits on your machine. Try the SEO Crawler or compare the plans.

Download Crawl Cove