When a client asks "what exactly do you test?", this is the answer, and the detail behind the On-Page SEO Checker landing page. Crawl Cove ships 49 deterministic checks across nine categories: HTTP, meta tags, content, links, security, images, an auto-detected WordPress pack, AI-readiness checks, and performance (Core Web Vitals).
Every check has a stable id, a semantic version, a default severity, and many expose tunable thresholds you can adjust per site in the Check Registry. Findings are keyed by (check id, URL), which is what lets the app compare runs and show honest deltas over time.
Note
Findings are deterministic, never guessed. Every row below is produced by a coded check function, not by an AI. The optional LLM layer only explains a finding in plain English; it can never invent, remove, or rescore one. When you put a number in front of a client, it came from a rule you can point to.
Why the version matters: when you improve a check, its semantic version changes, so run-to-run deltas stay honest rather than silently shifting because the rule moved under you. And because findings are keyed by (check id, URL), the same issue on the same page is tracked as one continuous thread across audits. That thread is the foundation of proving a fix worked.
Note
Accurate as of Crawl Cove 0.1.8 (September 2026). This reference is checked against the check registry in the shipping release. Crawl Cove is actively developed, so future updates are likely to add new checks and categories, and existing checks are sharpened between releases. Treat the count above as today's snapshot, not a permanent ceiling. Your installed version's Check Registry is always the authoritative, up-to-date list.
Note
What "full audits only" means. A few checks reach their conclusion from
something being absent, so they only run when the crawl saw everything it
found. Crawl Cove treats a crawl as partial if you scoped it with include or
exclude patterns, if Max pages (ships at 500) or Max depth (ships at 5)
turned an in-scope URL away, or if fewer than 90% of the URLs it discovered were
fetched. On a partial crawl broken-internal-link is withheld entirely, and
canonical drops only its "target not found" case while its missing and
malformed cases carry on: a link or canonical pointing at a page the crawl never
reached cannot be told apart from one pointing nowhere. orphan-page and
weakly-linked-page soften their severity instead of going quiet, and the two
sitemap checks have always needed a full audit. Every audit states on its own
screens how much of the site it covered, so a withheld check is never a silent
one.
HTTP
| Check id | Default severity | Detects | Tunable |
|---|---|---|---|
status-code |
Critical | Pages that failed to fetch or returned 4xx/5xx (redirects are out of scope). A 429 is the one exception: it is reported at Low, saying the page was not checked, because the server was throttling the crawler rather than serving a broken page. | none |
redirect-chain |
Low | Pages reached only after one or more redirects, showing the full hop chain. | min hops (1) |
Meta tags
| Check id | Default severity | Detects | Tunable |
|---|---|---|---|
title |
High | Missing/empty or wrong-length <title>. |
min 30 / max 60 chars |
duplicate-title |
High | Indexable pages sharing a title with other indexable pages. | none |
meta-description |
Medium | Missing/empty or wrong-length meta description. | min 70 / max 160 chars |
duplicate-meta-description |
Medium | Indexable pages sharing a meta description. | none |
canonical |
Medium | Missing or non-absolute canonical on indexable pages. Its third case, a same-origin canonical pointing at a URL the crawl never saw, is full audits only. | none |
canonical-to-noindex |
High | Canonical resolves to a page that is itself noindex or non-200. | none |
canonical-to-redirect |
Medium | Canonical points at a redirecting URL, not the redirect's final destination. | none |
insecure-canonical |
Medium | HTTPS page canonicalised to its http:// version. |
none |
social-meta |
Low | Indexable pages missing Open Graph / Twitter Card tags used for social and AI share previews. | none |
html-lang |
Low | Missing or invalid root <html lang> attribute. |
none |
hreflang |
Medium | Invalid hreflang codes, missing self-reference, missing return links, or alternates pointing at non-indexable URLs. | none |
robots-noindex |
High | Surfaces pages excluded from search by a robots directive (noindex, or its synonym none), from either the robots meta tag or the X-Robots-Tag response header, so the exclusion can be confirmed as intentional. |
none |
structured-data-missing |
Low | Substantive indexable pages with no JSON-LD structured data. | min 100 words |
url-consistency |
Low | Inconsistent casing, trailing slashes, or redundant query parameters that risk duplicate-content dilution. | none |
sitemap-non-indexable-url |
Medium | Sitemap URLs that are not canonical, indexable 200 pages (uncrawled, redirecting, non-200, or noindex). Full audits only. | none |
page-missing-from-sitemap |
Low | Indexable, self-canonical crawled pages not listed in the XML sitemap. Full audits only. | none |
Content
| Check id | Default severity | Detects | Tunable |
|---|---|---|---|
heading-h1 |
High | Not exactly one H1 on the page. | expected H1 count (1) |
heading-order |
Low | A skipped heading level (e.g. H2 → H4). Readability and accessibility, not ranking. | max heading level jump (1) |
thin-content |
Low | Indexable pages below the word-count floor; near-empty pages flagged at higher severity. | min 200 words, very-thin at 60 |
content-under-structured |
Low | Substantial pages (≥ 600 words) with fewer than 2 H2-H6 subheadings. | min words (600), min subheadings (2) |
near-duplicate-content |
Medium | Pages whose main body content is near-identical to other indexable pages (SimHash distance). | max fingerprint distance (10 bits) |
content-depends-on-js |
Medium | Critical content or structured data that only appears after JavaScript runs (raw-vs-rendered diff). | none |
schema-validity |
Medium | Malformed JSON-LD, or structured data missing required Schema.org properties. | none |
Links
| Check id | Default severity | Detects | Tunable |
|---|---|---|---|
broken-internal-link |
High | Internal links to same-site URLs not found in the crawl (typo, deleted page, wrong path). Full audits only. | none |
orphan-page |
Medium* | Indexable non-seed pages with zero internal inbound links. | none |
internal-link-to-redirect |
Medium | Internal links pointing at a redirect, not the final URL. | none |
excessive-crawl-depth |
Low | Indexable pages too many clicks from the seed. | max depth (4) |
nofollow-internal-link |
Low | Internal links marked rel="nofollow", which block internal PageRank flow. |
none |
dead-end-page |
Low | Indexable pages with no outbound internal links. | none |
weakly-linked-page |
Low* | Pages with some but fewer than the recommended inbound links (complements orphan-page). |
min inbound (2) |
internal-link-opportunity |
Info | Suggests missing contextual internal links to topically related pages (an opportunity, not a defect). | shared terms (2), relatedness (0.18), max/page (5), boilerplate cutoff (0.5) |
non-descriptive-anchor-text |
Low | Generic anchors ("click here", "read more"). | none |
excessive-internal-links |
Low | Pages whose internal link count exceeds the recommended maximum (sprawl). | max links (300) |
*Severity softens on partial/scoped crawls.
Note
The two checks marked with an asterisk, orphan-page and weakly-linked-page,
automatically soften their severity on partial or scoped crawls. Inbound
links may well exist on pages you chose not to crawl, so the app refuses to cry
"orphan" with full confidence when it can't see the whole link graph. Honesty
over false alarms.
Security
| Check id | Default severity | Detects | Tunable |
|---|---|---|---|
mixed-content |
High | HTTPS pages referencing insecure http:// resources in image srcs, link hrefs, or JSON-LD. |
none |
Images
| Check id | Default severity | Detects | Tunable |
|---|---|---|---|
image-alt |
Medium | Pages with images whose alt attribute is absent entirely, or present but only whitespace. A deliberate alt="" is correct markup for a decorative image and is not flagged. |
min images missing (1) |
WordPress (only on auto-detected WordPress sites)
These checks run only when the crawler auto-detects WordPress. A non-WordPress site gets zero WP probes and zero WP findings. See How the crawler behaves.
| Check id | Default severity | Detects | Tunable |
|---|---|---|---|
wp-generator-version-exposed |
Low | WordPress version leaked via the generator meta tag. | none |
wp-xmlrpc-reachable |
Medium | /xmlrpc.php publicly reachable (DDoS/brute-force vector). |
none |
wp-default-permalinks |
Low | Default ?p=N permalinks instead of descriptive URLs. |
none |
wp-uploads-listing-open |
Medium | /wp-content/uploads/ returns an open directory index. |
none |
wp-missing-image-dimensions |
Low | Images without width/height attributes (causes layout shift). | none |
AI-readiness
| Check id | Default severity | Detects | Tunable |
|---|---|---|---|
ai-bot-access |
Medium | AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) blocked in robots.txt. | none |
ai-llms-txt |
Low | No /llms.txt content-map file present (advisory). |
none |
ai-extractability |
Low | Pages structurally hard for AI engines to quote (names the missing signals). | none |
ai-author-authority |
Low | Article structured data missing an author or author identity link (the E-E-A-T signal generative engines use). | none |
ai-content-freshness |
Low | Article structured data missing a dateModified freshness signal. |
none |
ai-answer-structure |
Info | Substantial content pages with no question/answer structure for generative engines to extract. | min 400 words |
ai-entity-markup |
Low | No Organization/Person entity markup with sameAs links for entity disambiguation (advisory). |
none |
Performance
| Check id | Default severity | Detects | Tunable |
|---|---|---|---|
performance-core-web-vitals |
Medium | Pages whose Core Web Vitals are in Google's poor range, from CrUX field data (preferred) or Lighthouse lab data (fallback). | LCP (4000 ms), INP (500 ms), CLS (0.25) |
Tuning these to a client
Every threshold above is a per-site default, not a hard-coded law. An e-commerce site with faceted URLs needs different link rules than a blog. In the Check Registry you can enable or disable any check per site, override its severity, and adjust its thresholds, all respected by the audit pipeline on the next crawl. See Tuning checks with the Check Registry.