Skip to content
Meta Medium severity

Missing Canonical Tag: What It Costs You

An indexable page with no rel=canonical at all, or one that isn't a valid absolute URL, leaves a search engine to guess which version of the page is real.

A canonical tag is how you tell a search engine which URL should hold a page's ranking signals when more than one URL could serve the same content. Leave it out, or get it wrong, and the search engine decides for itself, which is rarely the outcome you would have chosen.

What this finding means

Crawl Cove's canonical check looks at every indexable page (HTTP 200, not itself noindex) and flags it for the first of three problems that applies: no rel="canonical" at all, a canonical value that does not parse as a valid absolute URL, or a same-origin canonical pointing at a URL the crawl never found anywhere on the site.

A real finding looks like this:

no rel=canonical link on an indexable page

or, for a malformed value:

canonical is not a valid absolute URL: "/blog/post-name"

or, for a same-origin target that does not exist in the crawl:

canonical "https://example.com/deleted-section" was not found among crawled in-site URLs

One finding per page, reporting the single most relevant issue rather than stacking several messages about the same tag.

Why it matters

Without a canonical, a search engine is left to work out on its own which of several similar URLs (with and without a trailing slash, with tracking parameters, printer-friendly versions) is the one that should rank. It usually gets this right often enough, but "usually" means your ranking signals can end up split across near-duplicate URLs instead of consolidated onto the one you actually want indexed, and you have handed over a decision that was cheap to make yourself.

How to fix it

  1. Add a <link rel="canonical" href="..."> to any indexable page that is missing one, pointing at the preferred, fully-qualified absolute URL for that content, most often the page itself.
  2. Fix a malformed canonical by making it absolute: include the scheme and host, not just a path.
  3. For a dangling same-origin canonical, confirm whether the target moved, was deleted, or was mistyped, then repoint the canonical at a URL that actually exists and is indexable.
  4. Re-crawl the site to confirm the finding clears.

False positives and edge cases

  • A same-origin canonical pointing at a URL the crawl never reached is not flagged on a partial crawl. A crawl scoped by include/exclude rules, or capped by a page or depth limit, has plenty of real, live pages it simply never got to; treating their absence as a broken canonical would be a false positive, so this branch only runs on a full, unscoped audit.
  • A canonical pointing at a URL outside a scoped crawl's own include/exclude patterns is never flagged, for the same reason: it was deliberately never a crawl candidate.
  • A canonical pointing at a page that resolves, but only via a redirect or to a noindex target, is not flagged here. Those are Canonical Pointing to a Redirect and Canonical Pointing to a Noindex Page, each with its own fix; this check only covers "missing," "malformed," or "target never seen in the crawl at all."
  • An off-site (cross-domain) canonical is validated only as an absolute URL, never checked against the crawl's own pages, since a legitimate cross-domain canonical (syndicated content pointing back at the original publisher) is not something this site's crawl could ever confirm either way.

Related reading

For a canonical that resolves but lands on a redirect instead of the final URL, see Canonical Pointing to a Redirect. For one that resolves to a page excluded from search entirely, see Canonical Pointing to a Noindex Page.

Frequently asked questions

Does every page need its own canonical tag?
Yes, at minimum a self-referential one pointing back at itself. That is the normal, correct pattern on the large majority of pages, and it costs nothing to add.
What counts as "not a valid absolute URL"?
A canonical value that fails to parse as a full http or https URL, most often a relative path (/page instead of https://example.com/page) or a typo that breaks the URL structure. Search engines can be inconsistent about resolving a relative canonical, so an absolute one removes the ambiguity entirely.
Does this check flag a canonical that points at a redirect or a noindex page?
No, those are reported separately as Canonical Pointing to a Redirect and Canonical Pointing to a Noindex Page, because each has its own, more specific fix. This check owns the broader "is there a valid canonical at all" question, and stays quiet once a canonical exists, parses correctly and points at a URL the crawl actually found.
Why is a dangling canonical target not always flagged?
On a crawl that only covers part of a site, whether scoped by include/exclude rules or cut short by a page or depth limit, a canonical can legitimately point at a real page the crawl simply never reached. Flagging that as broken would be a false positive, so this specific branch is skipped on a partial crawl.

Audit your site the easy way

Crawl Cove finds this on your machine, on every plan, and tells you exactly what to fix first. See the features or compare the plans.

Download Crawl Cove