Skip to content
Technical SEO 8 min read By The Crawl Cove team

Finding Orphan Pages: Crawler + Search Console

How to find orphan pages on your site using a crawler's link graph and Google Search Console together, why they build up, and how to fix each one.

Part of Technical SEO Guide: Crawl, Index, Speed

Key takeaways

  • An orphan page has no internal links pointing to it at all, so search engines struggle to find it and rarely revisit it once they do
  • A crawler's link graph finds every orphan reachable from your current site structure, and it needs to reach the whole site to be sure
  • Google Search Console's Internal Links report can surface pages a crawl already fixed or an old page a crawl no longer reaches, so the two sources answer different questions
  • Fix each orphan by linking it in deliberately, redirecting it to a live equivalent, or retiring it on purpose, never by accident

An orphan page is a page that exists and works fine if you visit it directly, but has no internal link pointing to it from anywhere else on your site. To find them, you need two different sources, because they catch two different kinds of orphan: crawl your site to build its current link graph, and check Google Search Console's Internal Links report for pages your crawl no longer reaches at all. This guide covers why orphan pages build up, how to find every one with each method, and what to do once you have the list.

What an orphan page actually is

A page is an orphan when nothing else on the site links to it: not the navigation, not the footer, not a related-content block, not another article's body text. The page itself might be well-written and technically sound. It loads fine, returns a normal 200 status, and is perfectly indexable. The only thing wrong with it is that a visitor, or a search engine, can only reach it by typing the exact URL or following an old external link.

That last point is why orphan pages are so easy to miss. Nothing about visiting the page tells you it is orphaned. There is no error, no warning, no broken layout. The only way to know is to look at the site's link graph from the outside and notice the page isn't in it.

Why orphan pages happen

Orphan pages are rarely deliberate. They build up for mundane, entirely ordinary reasons:

  • A landing page built for a campaign, linked only from an email or an ad, and never added to the site's own navigation or a related-content block once the campaign ended.
  • An older article dropped out of a category or archive listing after the section was redesigned, even though the article itself is still live.
  • A product or feature page whose parent listing page was rebuilt, and the individual page's link quietly didn't make the cut.
  • A page that used to sit in the main navigation, got removed from the menu during a redesign, and nobody added a replacement link anywhere else.

None of these show up as a build error or a broken link. The page keeps working; it just stops being reachable.

Why it costs you rankings

Search engines discover the overwhelming majority of pages by following links, not by guessing at URLs. A page with no internal links pointing to it is harder to discover in the first place, and once found, it is far less likely to be crawled again, because there is no fresh link signalling that it still matters. Internal links also carry ranking signal between pages on your own site; a page nothing links to receives none of that signal, no matter how good the content on it is.

There is no penalty for having an orphan page. The cost is entirely opportunity: a page that could be earning impressions and passing authority to and from related pages, sitting outside the graph doing neither.

Method 1: crawl the site and read its link graph

The most direct way to find orphan pages is to crawl the site the way a search engine does, then look for indexable pages nothing in that crawl pointed at.

Crawl Cove's orphan page finder builds your internal link graph as it crawls, then flags every indexable page that no other crawled page links to. Two details matter for trusting the result: a page linking to itself does not rescue it from the list, and the page you started the crawl from is never flagged, since a crawl seed has no "inbound link" to speak of by definition. That keeps the output a short list of genuine orphans instead of a pile of false positives from self-referential pages.

This method needs to reach the whole site to give a real verdict. If the crawl stops early, at a page cap or a depth limit, a page the crawl never got to look reachable or unreachable at random, so the check should be withheld rather than guessed at on a partial crawl. Raise the page limit until the crawl covers the site before trusting the orphan list.

Once you've connected Google Search Console, the orphan list gets more useful than a flat inventory: findings are ranked by the Search Console impressions each page already earns, so the orphan that's quietly still getting search impressions despite having no internal links surfaces above one nobody has ever seen. Without Search Console connected, the list falls back to ranking by severity alone.

The Crawl Cove Connections page showing a connected Google account, property picker, and last-sync status
Connect the account once, then pick a property per client. After that, orphan findings rank by real Search Console impressions instead of severity alone.

Tip

Run this after any redesign that touches navigation, category pages, or archive listings. Those are exactly the changes that create new orphans, and a crawl straight after the change catches them before they've had time to fall out of the index.

Method 2: check Google Search Console's Internal Links report

A crawl only ever sees the site as it exists today. Search Console remembers more than that: pages that were linked at some point in the past, picked up an external backlink, or were submitted in an old sitemap can stay indexed long after every internal link to them has disappeared. A page like that will not show up in a fresh crawl's link graph as "reachable but unlinked", because the crawl genuinely cannot reach it at all; it will only show up by comparing Search Console's records against what a crawl finds today.

In Search Console, open Links → Internal links and sort by link count, ascending. Pages with a very low or zero internal link count are your candidates. Cross-reference each one against the Pages report under Indexing: a page that is indexed despite having few or no internal links, according to Search Console's own record, is worth a look even if your current crawl doesn't reach it at all.

The two sources answer genuinely different questions. A crawl tells you what's unreachable right now, from the site as it exists today. Search Console tells you what Google still remembers, which can include pages your site has already structurally cut off. Use both; neither alone gives the complete picture.

Fixing each orphan page

Once you have a list from either or both sources, every orphan resolves to one of three deliberate outcomes.

Link it in, if it's worth keeping. Add a contextual internal link from at least one genuinely relevant page: a related article, a category or hub page, a closing "you might also like" block. A link from body content, next to relevant text, does more work than a generic footer link.

Redirect it, if a better page already covers the topic. If the orphan has been superseded by a newer or merged page, a 301 redirect to that equivalent passes on whatever link equity and search visibility the old page had accumulated, rather than leaving both pages competing for nothing.

Retire it on purpose, if it no longer earns its place. Not every orphan deserves rescuing. If the page is genuinely outdated and nothing else should link to it, let it go with an intentional 410 rather than leaving it to linger unlinked and unindexed by accident. The difference between an orphan and a deliberately retired page is that one happened to you and the other is a decision you made.

Keeping orphan pages from building back up

Fixing today's list matters less than stopping tomorrow's list from being the same size. A couple of habits keep it that way:

  • Link new pages the day they're published. Every new page should get at least one contextual link from an existing relevant page before it goes live, not sometime later "when there's time".
  • Re-check after every navigation or category redesign. These are the changes most likely to quietly cut a page's only inbound link. A crawl straight after the redesign catches it early.
  • Re-run the crawl after each fix. With Crawl Cove's versioned audits, you can compare the new crawl against the old one and confirm a page you linked, redirected, or retired actually left the orphan list, rather than assuming the fix landed.

If you're maintaining this across several client sites, doing the crawl-plus-Search-Console check on a repeatable schedule per project is far cheaper than rediscovering the same orphans months later. See pricing for how that scales across client projects.

Wrap-up

An orphan page is any indexable page nothing on your site links to, and it happens quietly through completely ordinary changes: a redesigned category page, a retired campaign, a menu item removed without a replacement. A crawler's link graph finds every orphan your site's current structure has created; Google Search Console's Internal Links report finds the older ones your structure has already forgotten about. Use both, then link, redirect, or retire each one on purpose. Re-run the crawl afterwards to confirm the fix actually held.

Frequently asked questions

What exactly counts as an orphan page?
A page that is live, indexable and returns 200, but has no internal link pointing to it from anywhere else on the site. A link from the page to itself does not count, and neither does an external backlink; the page has to be unreachable by following links from your own site.
Can Google still know about a page with no internal links?
Yes. If the page was linked to in the past, submitted in a sitemap, or has an external backlink, Google may have indexed it before the links disappeared and can keep it in the index for a long time afterwards. That is exactly why a crawl alone will not catch every orphan; Search Console remembers pages your current site structure has already forgotten.
Do orphan pages get penalised by Google?
No, there is no penalty. The cost is opportunity, not punishment: a page nothing links to is harder to discover, gets crawled and re-crawled less often, and receives none of the internal link equity that helps a page rank.

Audit your site the easy way

Crawl Cove finds these issues on your machine. Try the Orphan Page Finder and the Internal Link Checker or see every feature.

Download Crawl Cove