Sitewide Duplicate Content Finder

Last updated:

Go beyond a two-page comparison. Crawl a public site or paste a URL inventory to uncover exact copies, near-duplicate clusters, repeated templates, canonical conflicts, and the pages worth consolidating.

Breadth-first, same-host crawl. It follows ordinary HTML links and stops at the selected limit or the server time limit.

Privacy: only submitted public URLs and their public HTML are processed; there is no signup or report-storage step. Do not submit authenticated or private pages.

How This Sitewide Duplicate Content Audit Works

  1. Build a bounded inventory. Start a same-host link crawl or paste the important URLs from a sitemap, CMS, analytics, Search Console, or another crawler. Inventory mode is the safer choice for orphaned pages.
  2. Extract search signals. For each server-rendered HTML response, the audit records the final URL, status, title, first H1, canonical, robots directives, indexability, and readable content blocks.
  3. Measure template repetition. Navigation, headers, footers, scripts, forms, and similar chrome are removed first. Substantive blocks repeated across at least 40% of the successful crawl sample and at least three pages are measured as boilerplate and removed from a second, core-text representation.
  4. Compare both content views. The complete extracted main text prevents genuinely copied sections from disappearing merely because they recur across several pages; the de-boilerplated core view reduces template noise. Exact fingerprints and three-word shingle Jaccard and containment scores are calculated on both, and a strong result from either view can surface a pair.
  5. Cluster and prioritize. Qualifying pairs form transparent single-link content families. Each cluster reports direct-pair coverage and its weakest qualifying overlap, proposes a preferred-page candidate, and receives a context-aware consolidation or differentiation recommendation.
  6. Verify before changing URLs. Filter the full inventory, inspect canonical and indexing signals, export CSV, then add backlinks, traffic, conversions, and business value before making a final decision.

Why Sitewide Context Is Better Than a Pairwise Score

Two pages can look similar for very different reasons. Product variants may deliberately share specifications. Location pages may repeat service information but still need meaningful local evidence. Tracking parameters may produce exact URL copies. A sitewide audit exposes the pattern, surrounding signals, and scale.

Google describes redirects and rel="canonical" as strong canonicalization signals, while sitemap inclusion is weaker. Signals can reinforce one another, but Google still chooses the representative URL it considers best. That is why this tool reports canonical alignment rather than pretending a similarity percentage determines indexation. See Google’s canonicalization guidance.

  • Consolidate signals: retire obsolete copies and point links, sitemaps, canonicals, and redirects at one useful destination.
  • Protect distinct intent: keep pages that answer meaningfully different needs, then make their headings, evidence, copy, and internal links genuinely distinct.
  • Expose template risk: repeated legal text or product chrome may dominate thin pages even when the URLs are not literal copies.
  • Catch conflicting directives: near-identical pages with self-canonicals, cross-canonicals, noindex, or non-200 responses deserve different fixes.

How to Choose: Redirect, Canonical, Noindex, or Rewrite

Permanent redirect

Use a direct 301 or 308 when an old page has been replaced and visitors should always land on the stronger equivalent. Update internal links and sitemap entries too.

Canonical to the preferred version

Use a canonical for intentionally accessible alternate versions, such as some product variants or parameter views. The canonical destination should be indexable, return 200, and normally self-canonicalize.

Noindex

Use noindex when a page remains useful to people but should not appear in search, such as an internal result or account utility. Google advises against using robots.txt to control canonical selection because a blocked duplicate cannot be crawled to see its signals.

Keep and differentiate

When both pages serve distinct intent, improve more than wording. Add page-specific evidence, examples, products, local details, questions, media, and internal-link context. There is no safe universal “percentage unique” target.

Manual review

Escalate pages with little text, heavy templates, conflicting canonicals, pagination, translations, faceted navigation, or valuable historic signals. A crawler cannot infer business importance.

Limits You Should Know Before Acting

This is a focused free audit, not an unlimited rendering crawler. It reviews at most 50 public, same-host HTML pages and stops after roughly 42 seconds. It does not render JavaScript, log in, read private Search Console data, inspect backlinks, or know which page earns revenue. Similarity thresholds surface candidates; they do not diagnose a penalty.

For the best coverage, combine a link crawl with a separate URL inventory drawn from XML sitemaps, CMS exports, analytics landing pages, Search Console pages, server logs, and backlink tools. Run the same sample before and after consolidation so changes are comparable.

Next steps

Sitewide Duplicate Content Finder related tools and articles

Continue with the closest follow-up checks and guides based on this tool's topic, crawl intent, and optimization workflow.

Sitewide Duplicate Content Finder: FAQ

How does the sitewide duplicate content finder calculate similarity?
The audit extracts readable main-content blocks and compares two representations: complete extracted main text, plus core text with blocks repeated across at least 40% of the successfully crawled sample and at least three pages removed. Both use normalized fingerprints and overlapping three-word shingles. A pair qualifies when either representation clears the published threshold, while the repeated-word share stays visible for review.
What counts as an exact or near duplicate?
Exact means the normalized fingerprints match in either the complete extracted text or the de-boilerplated core text. Near-duplicate candidates are surfaced when either comparison reaches at least 58% Jaccard similarity, or at least 82% containment for pages with 80 or more words in that representation. These are transparent review thresholds, not Google penalties or universal SEO rules.
Does the tool render JavaScript?
No. It reviews the server-returned HTML. Content that appears only after client-side JavaScript runs may be missing, so a low word count or unexpected result on a JavaScript-heavy site should be verified with a rendering crawler.
How is the preferred page in a cluster selected?
The tool uses a deterministic triage score that favors a successful 200 response, indexability, a self-referencing or absent canonical, more core content, less repeated boilerplate, and a slightly shorter URL. It cannot see conversions, backlinks, revenue, or historical rankings, so the preferred page is a review candidate rather than an automatic decision.
Should every duplicate page be redirected?
No. Redirect obsolete URLs when users no longer need separate versions. Use a canonical when alternate versions must remain accessible but should consolidate search signals. Keep both pages when they serve distinct intent, and improve the unique value of each. Noindex can suit utility pages, but it is not a substitute for canonicalization.
What does boilerplate risk mean?
It is the share of extracted words contained in substantive blocks that recur across much of this crawl sample. High repetition can come from legitimate templates, product specifications, legal copy, calls to action, or extraction limitations. It is a diagnostic clue, not an error by itself.
How are pages grouped into duplicate clusters?
Qualifying page pairs are connected into single-link clusters, so A can be grouped with C when both match B even if A and C do not clear a threshold directly. The report shows direct-pair coverage and the weakest qualifying overlap—not the strongest edge—as a conservative coherence signal. Review every member before applying one action to an entire cluster.
Why might the crawl miss pages?
Link crawl mode can only discover URLs reachable from the start page in the bounded sample. It stops at 50 pages and about 42 seconds, skips non-HTML assets, remains on one hostname, and does not obey forms or JavaScript navigation. Use inventory mode to supply important URLs from a sitemap, CMS export, analytics, or Search Console.
Is my content stored?
Public URLs are sent to the Web Aloha endpoint so it can fetch and compare their returned HTML. The endpoint has no database or result cache and does not intentionally persist page content or reports. CSV exports are generated in your browser. Do not submit private, authenticated, preview, or confidential URLs.

Free 48-Hour Website Audit

Not sure what to fix first on your own website? We'll review it and tell you, in plain English. Free & non-obligatory.

Turn Duplicate Clusters into a Clear SEO Plan

We combine crawl evidence with rankings, links, conversions, and business value to decide what to merge, canonicalize, rewrite, or keep.