Free XML Sitemap Generator

Last updated:

Crawl up to 500 public URLs or build a sitemap from a URL list. Review redirects and exclusion reasons, then validate and export standards-based XML.

How the XML Sitemap Generator Works

This tool creates a valid XML sitemap file for your website in two ways:

  1. Crawl mode, enter a website URL and this sitemap crawler checks up to 500 same-host URLs discovered in static HTML links. The report shows included, excluded, and failed URLs with the response and indexing signals it could inspect.
  2. Choose a query policy, the safe default drops all query strings before crawling. Keep non-tracking parameters only when those variants are intentional canonical pages; common campaign parameters are always removed.
  3. Manual mode, paste a list when you know exactly which pages belong in the sitemap. This is also the route for JavaScript-only navigation and known orphan URLs, but manual URLs are normalized rather than fetched or audited.
  4. Add only verified metadata, priority and changefreq are optional and ignored by Google. Include lastmod only when it reflects a real, meaningful update to the page.
  5. Review and edit, after generation, review the URL list. Uncheck any pages you want to exclude. The XML preview updates in real time.
  6. Validate and download, the browser checks the generated XML structure, URL count, and 50 MB uncompressed protocol ceiling before copy or download. Publish it on your site, reference it in robots.txt, and optionally submit it in Google Search Console.

Why XML Sitemaps Matter for SEO

An XML sitemap supplies canonical URL candidates and optional, accurate modification dates in a machine-readable format:

  • Discovery support, a sitemap gives search engines a direct list of URLs you prefer them to crawl and potentially show. It is especially useful when a site is new, large, media-rich, or not comprehensively linked.
  • Canonical consistency, listing only absolute, preferred URLs reinforces the canonical version you want search engines to consider.
  • Useful monitoring, Search Console can report when Google read a submitted sitemap and flag processing problems. Submission remains a hint, not a guarantee of crawling or indexing.
  • Accurate freshness, Google may use trustworthy lastmod values when they consistently reflect significant page changes. The crawler cannot derive that truth, so this generator leaves the field off by default.

After creating your sitemap, validate it with our Sitemap Checker & Validator to make sure everything is technically correct. Also consider configuring your robots.txt to reference the sitemap with a Sitemap: directive. Use the Internal Link Analyzer to inspect links, but remember that no link crawler can prove a URL is orphaned without another complete source of URLs to compare. For multilingual annotations, use the Hreflang Tag Generator.

XML Sitemap Best Practices

Follow these guidelines to get the most out of your XML sitemap:

  • Only include canonical, indexable pages, exclude noindex pages, redirects, 404s, paginated archives, and duplicate URLs. Every URL in your sitemap should return a 200 status and be the canonical version.
  • Keep lastmod accurate, only update the lastmod date when the page content actually changes. Fake or auto-updating lastmod dates erode trust with search engines and may cause them to ignore the field entirely.
  • Respect the actual protocol limits, one sitemap may contain up to 50,000 URLs and be up to 50 MB uncompressed. Split it only when it exceeds either limit; a sitemap index can reference the resulting files. The generator’s 500 checked-URL ceiling is unrelated.
  • Reference it in robots.txt, add Sitemap: https://yourdomain.com/sitemap.xml to your robots.txt file. This helps crawlers find your sitemap even if they have never seen your site before.
  • Submit to Search Console, after publishing, submit the sitemap through Google Search Console to monitor access and processing. Treat submission as a discovery hint, not evidence that every URL was crawled or indexed.
  • Automate when possible, most CMS platforms (WordPress with Yoast/Rank Math, Shopify, Astro, Next.js) can generate and update sitemaps automatically. Use manual sitemap generation primarily for static sites or custom setups.

For a deeper understanding of sitemap discovery and technical SEO, explore our guide to sitemaps, and our SEO services.

Next steps

XML Sitemap Generator related tools and articles

Continue with the closest follow-up checks and guides based on this tool's topic, crawl intent, and optimization workflow.

Sitemap Generator: FAQ

How does Crawl Website discover pages?
The sitemap crawler starts at the entered URL and follows same-host links found in static HTML anchor tags. It checks up to four URLs concurrently and reports each checked URL, final URL, status, redirect, content type, canonical, meta robots, and X-Robots-Tag where available.
Which pages does crawl mode leave out?
It excludes non-200 responses, non-HTML files, meta robots or X-Robots-Tag noindex pages, cross-canonical pages, duplicate final URLs, and cross-host redirects. The crawl report gives a reason for every checked URL instead of silently dropping it.
What important crawl limitations should I know?
This is a static-HTML crawl. It does not execute JavaScript, authenticate, evaluate robots.txt, or discover true orphan URLs that are not linked from crawled HTML. It checks at most 500 discovered URLs or runs for 45 seconds, and each request can take up to ten seconds.
How are query parameters handled?
By default crawl mode removes every query string before queueing a URL, which avoids session, filter, calendar, and faceted-navigation crawl traps. The opt-in keep mode still removes common tracking parameters and sorts the remaining parameters deterministically. Use it only when query variants are distinct canonical pages.
Why were fewer pages found than expected?
Navigation may require JavaScript, the start page may not link to deeper pages, robots or a firewall may block this crawler, or the crawl may hit its URL or time limit. True orphan URLs cannot be found by link crawling; paste known URLs in Manual Entry and compare against CMS exports, analytics, or Search Console.
What validation does Manual Entry perform?
It accepts syntactically valid HTTP or HTTPS URLs, adds https:// to a bare hostname, removes fragments and common tracking parameters, sorts remaining parameters, and deduplicates normalized URLs. It also enforces one host and the 2,048-character loc limit. It does not fetch URLs or verify public reachability, status, canonical, indexability, ownership, or robots.txt.
How should I use priority, changefreq, and lastmod?
All three are optional and default to none. Google ignores priority and changefreq. Select today for lastmod only when every included page was meaningfully updated today. In crawl mode the server does not derive real modification dates from the pages.
Why can the XML preview look incomplete?
The on-page preview shows only the first 100 lines for performance, but Copy XML and Download sitemap.xml use the full generated document for all selected URLs. Uncheck unwanted rows and verify the URL total before exporting.
Is 500 URLs the XML sitemap limit?
No. It is this free crawler’s capacity. The sitemap protocol permits up to 50,000 URLs or 50 MB uncompressed in one sitemap. Above either protocol limit, split the output and optionally use a sitemap index.
Does submitting a sitemap guarantee indexing?
No. Uploading a sitemap, listing it in robots.txt, or submitting it in Google Search Console helps discovery and monitoring, but submission is only a hint and does not guarantee crawling or indexing.
What data is sent or stored in each mode?
Manual Entry, selection, XML generation, copying, and downloading run in your browser. Crawl mode sends the start URL, URL limit, and query policy to the Web Aloha server, which requests public pages and returns the crawl report. The endpoint does not cache results or write them to a database.

Free 48-Hour Website Audit

Not sure what to fix first on your own website? We'll review it and tell you, in plain English. Free & non-obligatory.

Need Help with Technical SEO?

We help businesses set up sitemaps, robots.txt, crawl directives, and full technical SEO foundations.