Skip to content
Technical SEO 6 min read By The Crawl Cove team

How to Catch Canonical Tag Template Bugs

A canonical bug in a shared template does not break one page, it breaks every one it renders, including one we shipped and thought we'd fixed first try.

Part of Technical SEO Guide: Crawl, Index, Speed

Key takeaways

  • A canonical tag mistake in a shared template affects every page that template renders, not one page, so it is worth catching before it ships rather than after a crawl finds it
  • A fix that only changes the visible canonical string can still leave every other generated URL on the page, RSS alternates, asset links, sitemap references, pointing at the wrong root
  • We found and fixed exactly this on our own site: a host-level routing quirk made crawlcove.com serve duplicate URLs under /index.php, each with a canonical that leaked the docroot and itself redirected
  • Testing the fix against the request shape that broke it is not enough; test the variants nobody thought to write a rule for, because a rewrite rule is only as good as the request shapes you actually sent it

A single mistyped canonical tag is a one-page problem. A canonical tag bug in a shared template, a layout, a controller, or a URL-generation helper, is a whole-site problem wearing a one-page disguise, because every page rendered through that template inherits the same wrong value. We found one of these on our own site, fixed it twice before it was actually fixed, and are writing down what it took to catch properly.

The bug: a front controller that shouldn't have answered

crawlcove.com's host maps incoming requests into public_html/ before .htaccess gets a chance to run. That meant any request naming the front controller directly, /index.php, and via PATH_INFO, /index.php/<any real route>, served a full 200 response: an unbounded space of duplicate URLs for every real page on the site.

The part that made it a genuine SEO problem, not just an ugly URL, was the canonical tag those duplicate pages carried. Every one of them declared:

<link rel="canonical" href="https://crawlcove.com/public_html/index.php">

That single value is wrong twice over. It leaks the server's docroot path into a public-facing tag, and the URL it points to is not a final page at all, it 301s. A canonical pointing at a redirect is close cousin to a canonical pointing at a noindex page: in both cases, the tag is telling search engines to consolidate signal onto a URL that cannot itself be indexed, which defeats the point of having a canonical at all.

Why it happened

Laravel resolves its base URL from SCRIPT_NAME via Symfony's prepareBaseUrl(). On a host that routes requests through the front controller before its own rewrite rules apply, SCRIPT_NAME reads as /public_html/index.php, so that string ends up baked into every URL the framework generates for that request, not just the canonical.

The fix that looked complete and wasn't

The first patch normalised the canonical tag directly in the page layout. Tested against the exact request shape that had exposed the bug, SCRIPT_NAME=/public_html/index.php, the canonical came back clean. It would have been easy to ship that and move on.

It wasn't actually fixed. The same test that confirmed the canonical also showed the RSS feed alternates, font preloads, and every asset URL on the page still carrying /public_html/index.php/ in front of them. A search engine crawling that "fixed" page would read a correct canonical tag and then follow a page full of links straight back into the duplicate URL space the canonical was supposed to consolidate away from.

The canonical tag was the most visible symptom, not the actual defect. The real fix had to sit underneath it, at the point where the framework decides what its own root URL is (URL::forceRootUrl() in our case), so every generated URL on every page inherits the correct value at once. A narrow fix to the one tag a test happened to check is a trap worth naming: if a bug can put the wrong root into one URL, check whether it put the wrong root into all of them before declaring it fixed.

Other ways a canonical template bug shows up

Our incident was a host-routing quirk specific to one server configuration, but the general shape recurs in different forms:

  • A hardcoded absolute canonical copied into a shared partial. A header or layout component with a literal canonical URL, copied from one page and never made dynamic, silently canonicalises every page that includes it onto the one page it was copied from.
  • A reverse proxy or CDN rewriting the Host header. If the app computes its base URL from the incoming Host header and a proxy in front of it changes that header, every canonical the app generates can point at the wrong domain, most commonly an internal or staging hostname leaking into production markup.
  • Trailing slash or case inconsistency in a URL helper. A route helper that is inconsistent about a trailing slash generates a canonical that does not match the URL the page is actually served at, which some search engines then read as "this page disagrees with itself about its own address."
  • Over-canonicalization on paginated templates. A shared pagination partial that always points its canonical at page one, rather than self-referencing each page, tells search engines that pages two onward do not need to exist independently, even when they contain genuinely different content.

Each of these shares the same shape as our own bug: one shared piece of code, many affected pages, and a fix that can look complete after checking only the symptom that got reported.

How to actually catch these before they ship

Check canonical values across every template, not just the one you changed. A canonical fix scoped to the page type where a bug was found can miss the same defect sitting in a different template that happens to share the same helper. Duplicate without user-selected canonical covers what it looks like from Google's side when a canonical fails to do its job, including two live examples we found on our own site while writing that post.

Test request shapes nobody deliberately wrote a rule for. After our first .htaccess rewrite rule deployed, /index.phpx, /index.php5, and /index.phpfoo still served the homepage: the host was matching anything merely starting with index.php, which no amount of reading the rule would have revealed. We only found it by manually trying variants the rule wasn't written to anticipate. A rewrite rule, or any pattern-based fix, is only as correct as the request shapes you actually sent it.

Re-crawl after the fix, not just the one URL that reported it. A missing canonical tag and a wrong one fail the same way from a search engine's perspective: both leave it to guess. Crawling the whole site again after a canonical fix, and comparing the result against the crawl that found the bug, is the only way to confirm the fix reached every template it needed to, not just the page where someone happened to notice.

Wrap-up

A canonical template bug is not one mistake, it is one mistake multiplied by every page a shared piece of code touches. Ours leaked a docroot path and pointed at a redirect, across an unbounded space of duplicate URLs, and our first attempt at fixing it addressed the symptom a test happened to check rather than the underlying URL resolution. Catching the real thing meant checking every generated URL on the page, not just the canonical tag, and testing request shapes deliberately chosen to break an assumption rather than confirm one. Crawl Cove's own crawler is how we found the second, wider version of the bug: a full re-crawl after the "fix" still showed the old root in the page's other URLs, which is exactly the kind of regression a one-URL spot check cannot catch.

Frequently asked questions

What is a canonical tag template bug?
It is a canonical tag mistake baked into a shared layout, controller, or URL-generation helper rather than a single page, so every page rendered through that template inherits the same wrong value. A one-off typo affects one URL; a template bug affects a whole class of them at once, which is why it is worth catching in review rather than one crawl report at a time.
What actually happened on your own site?
crawlcove.com's host was mapping incoming requests into the public_html directory before .htaccess ran, so any request naming the front controller, "/index.php" and, via PATH_INFO, "/index.php/any-real-route", served a 200 response. Every one of those pages carried a canonical of "https://crawlcove.com/public_html/index.php", which leaked our docroot path and was itself a URL that redirected, so the canonical pointed at a redirect rather than a final page.
Why didn't fixing the canonical tag actually fix it?
Our first patch normalised the canonical string in the page layout, and a test built around that one request shape showed a clean canonical. But every other generated URL on the page, RSS feed alternates, font preloads, asset links, still carried the same wrong root, because the real fault was the framework's base URL resolution, not the one line of markup that happened to be the most visible symptom of it. A crawler reading that "fixed" page still followed working links back into the duplicate space.
How do you actually catch these before they ship?
Crawl the site the way a search engine does and check canonical values across every template your site uses, not just the page type you changed. Then test request shapes nobody wrote a rule for on purpose: trailing slashes, alternate capitalisation, near-miss paths. Our own front-controller rewrite rule looked correct until we manually tried "/index.phpx" and "/index.php5" and found the host was matching anything merely starting with "index.php", not the exact string the rule assumed.

Audit your site the easy way

Crawl Cove finds these issues on your machine. Try the Duplicate Content Checker and the SEO Crawler or see every feature.

Download Crawl Cove