Skip to content
Content Medium severity

Near-Duplicate Content: Copycat Pages

Two or more pages whose body text is near-identical split ranking signal between them, even when their titles and URLs are completely different.

Two pages do not have to look the same to compete with each other. A product page and a near-copy of it published under a different URL, or a location page templated from the same paragraph with two words swapped, read as different pages to a visitor and as the same page to a search engine trying to decide which one to rank.

What this finding means

Crawl Cove's near-duplicate-content check compares the main body text of every indexable page (HTTP 200, not noindex) against every other indexable page, using a content fingerprint that tolerates small wording changes rather than requiring an exact match. Two pages within a small distance of each other on that fingerprint are near-duplicates. Each affected page gets one finding, naming its closest counterpart, how similar it is, and how many other near-duplicates it has.

A real finding looks like this:

near-duplicate content: ~91% similar to 2 other indexable page(s): https://example.com/page-b, https://example.com/page-c

The list is capped at the ten closest counterparts, with an honest "+N more" if a page belongs to a larger cluster, and the full list of affected URLs is always available in the findings drawer even when the evidence sentence is capped.

Why it matters

When a search engine finds two or more pages that say almost the same thing, it does not rank them both, it picks one and largely ignores the rest, splitting whatever links and relevance signal that content earned across pages that were competing for the same result rather than reinforcing each other. This is a common side effect of templated content: location pages built from one paragraph with the city name swapped, product variants that differ only in size or colour, or a blog post duplicated to target a slightly different keyword. Each new near-duplicate page dilutes the group rather than adding to it.

How to fix it

  1. Open the page named in the finding and its listed counterparts to confirm how much of the body text genuinely overlaps.
  2. Consolidate into one page with a 301 redirect from the others, if the pages serve the same intent and do not need to exist separately.
  3. Add a rel="canonical" pointing at the preferred version, if you need to keep the URLs live for another reason (tracking, a paid campaign, and similar).
  4. Differentiate the content if the pages genuinely serve different intents and should both rank, by rewriting the shared paragraphs rather than templating them.
  5. Re-crawl the site to confirm the fingerprint distance has widened and the finding clears.

False positives and edge cases

  • A page with a canonical tag to another URL is never flagged. It already declares itself a known duplicate, so a separate finding here would be noise.
  • Paginated archive pages are excluded entirely, since a page 2 sharing most of its template with page 1 and page 3 is expected, not a defect.
  • This compares body text, not titles or URLs. A page can have a completely unique title and still be flagged here if its paragraphs are near-identical to another page's; see Duplicate Title Tags for the separate, title-only version of this problem.
  • A handful of reworded sentences is not enough to clear this check. The similarity threshold tolerates minor wording changes on purpose, so a templated page with two words swapped still matches; genuinely rewriting the shared sections is what clears it.

Related reading

For the canonical-tag side of consolidating duplicate pages, see Canonical Points to a Redirect. For the title-only version of duplicate content, see Fix Duplicate Title Tags and Meta Descriptions.

Frequently asked questions

How similar do two pages have to be before this fires?
Roughly 84% bit-similar or closer, measured with a content fingerprint (SimHash) that tolerates a reworded sentence or two but not a genuinely different page. That threshold is deliberately loose enough to catch templated or lightly-reworded duplicates, and tight enough that two pages covering the same topic in their own words do not collide.
Is this the same as duplicate title tags?
No, and a page can fail one check without the other. Duplicate titles compare the `<title>` element; this check compares the page's main body text directly, so two pages with completely different titles can still be flagged here if the paragraphs underneath are near-identical, and two pages with the same title can pass this check if their content genuinely differs.
Does a page with a canonical tag pointing elsewhere still get flagged?
No. A page that already canonicalises to another URL has already declared itself a known duplicate, so flagging it here would just repeat that fact rather than surface a new problem.
What about paginated archive pages, which naturally share most of their content?
Also excluded. A page 2 of a blog archive is expected to share the bulk of its template and navigation with page 1 and page 3, so paginated series are left out of the comparison entirely rather than producing noise.

Audit your site the easy way

Crawl Cove finds this on your machine, on every plan, and tells you exactly what to fix first. See the features or compare the plans.

Download Crawl Cove