Sitewide Duplicate Content Finder
Last updated:
Go beyond a two-page comparison. Crawl a public site or paste a URL inventory to uncover exact copies, near-duplicate clusters, repeated templates, canonical conflicts, and the pages worth consolidating.
Breadth-first, same-host crawl. It follows ordinary HTML links and stops at the selected limit or the server time limit.
One same-host URL per line. A first CSV column containing URLs also works. The first 50 unique valid URLs are audited.
Privacy: only submitted public URLs and their public HTML are processed; there is no signup or report-storage step. Do not submit authenticated or private pages.
Priority review
Duplicate clusters
Page inventory
Signals for every fetched URL
| Recommended review |
|---|
Found issues? That’s exactly what our SEO audit fixes.
Our SEO Audit + Essentials package ($555, one-time) finds and fixes exactly these kinds of problems, with a prioritized plan in plain English.
Looks clean here. How is the rest of your SEO?
One passing check doesn’t make a ranking strategy. We audit the full picture: content, technical SEO, internal links, and search intent.
Want us to act on this result? Fix these findings →
Prefer the partnership route? Refer clients to us, earn 10% →
How This Sitewide Duplicate Content Audit Works
- Build a bounded inventory. Start a same-host link crawl or paste the important URLs from a sitemap, CMS, analytics, Search Console, or another crawler. Inventory mode is the safer choice for orphaned pages.
- Extract search signals. For each server-rendered HTML response, the audit records the final URL, status, title, first H1, canonical, robots directives, indexability, and readable content blocks.
- Measure template repetition. Navigation, headers, footers, scripts, forms, and similar chrome are removed first. Substantive blocks repeated across at least 40% of the successful crawl sample and at least three pages are measured as boilerplate and removed from a second, core-text representation.
- Compare both content views. The complete extracted main text prevents genuinely copied sections from disappearing merely because they recur across several pages; the de-boilerplated core view reduces template noise. Exact fingerprints and three-word shingle Jaccard and containment scores are calculated on both, and a strong result from either view can surface a pair.
- Cluster and prioritize. Qualifying pairs form transparent single-link content families. Each cluster reports direct-pair coverage and its weakest qualifying overlap, proposes a preferred-page candidate, and receives a context-aware consolidation or differentiation recommendation.
- Verify before changing URLs. Filter the full inventory, inspect canonical and indexing signals, export CSV, then add backlinks, traffic, conversions, and business value before making a final decision.
Why Sitewide Context Is Better Than a Pairwise Score
Two pages can look similar for very different reasons. Product variants may deliberately share specifications. Location pages may repeat service information but still need meaningful local evidence. Tracking parameters may produce exact URL copies. A sitewide audit exposes the pattern, surrounding signals, and scale.
Google describes redirects and rel="canonical" as strong canonicalization signals, while sitemap inclusion is weaker. Signals can reinforce one another, but Google still chooses the representative URL it considers best. That is why this tool reports canonical alignment rather than pretending a similarity percentage determines indexation. See Google’s canonicalization guidance.
- Consolidate signals: retire obsolete copies and point links, sitemaps, canonicals, and redirects at one useful destination.
- Protect distinct intent: keep pages that answer meaningfully different needs, then make their headings, evidence, copy, and internal links genuinely distinct.
- Expose template risk: repeated legal text or product chrome may dominate thin pages even when the URLs are not literal copies.
- Catch conflicting directives: near-identical pages with self-canonicals, cross-canonicals, noindex, or non-200 responses deserve different fixes.
How to Choose: Redirect, Canonical, Noindex, or Rewrite
Permanent redirect
Use a direct 301 or 308 when an old page has been replaced and visitors should always land on the stronger equivalent. Update internal links and sitemap entries too.
Canonical to the preferred version
Use a canonical for intentionally accessible alternate versions, such as some product variants or parameter views. The canonical destination should be indexable, return 200, and normally self-canonicalize.
Noindex
Use noindex when a page remains useful to people but should not appear in search, such as an internal result or account utility. Google advises against using robots.txt to control canonical selection because a blocked duplicate cannot be crawled to see its signals.
Keep and differentiate
When both pages serve distinct intent, improve more than wording. Add page-specific evidence, examples, products, local details, questions, media, and internal-link context. There is no safe universal “percentage unique” target.
Manual review
Escalate pages with little text, heavy templates, conflicting canonicals, pagination, translations, faceted navigation, or valuable historic signals. A crawler cannot infer business importance.
Limits You Should Know Before Acting
This is a focused free audit, not an unlimited rendering crawler. It reviews at most 50 public, same-host HTML pages and stops after roughly 42 seconds. It does not render JavaScript, log in, read private Search Console data, inspect backlinks, or know which page earns revenue. Similarity thresholds surface candidates; they do not diagnose a penalty.
For the best coverage, combine a link crawl with a separate URL inventory drawn from XML sitemaps, CMS exports, analytics landing pages, Search Console pages, server logs, and backlink tools. Run the same sample before and after consolidation so changes are comparable.
Next steps
Sitewide Duplicate Content Finder related tools and articles
Continue with the closest follow-up checks and guides based on this tool's topic, crawl intent, and optimization workflow.
Sitewide Duplicate Content Finder: FAQ
How does the sitewide duplicate content finder calculate similarity?
What counts as an exact or near duplicate?
Does the tool render JavaScript?
How is the preferred page in a cluster selected?
Should every duplicate page be redirected?
What does boilerplate risk mean?
How are pages grouped into duplicate clusters?
Why might the crawl miss pages?
Is my content stored?
Free 48-Hour Website Audit
Not sure what to fix first on your own website? We'll review it and tell you, in plain English. Free & non-obligatory.
Turn Duplicate Clusters into a Clear SEO Plan
We combine crawl evidence with rankings, links, conversions, and business value to decide what to merge, canonicalize, rewrite, or keep.