Robots.txt Tester, Validator & Path Simulator
Last updated:
Fetch a public robots.txt file or paste one for browser-only analysis. Review line-specific findings, edit and copy the file, then run a practical user-agent and URL-path diagnostic.
Fetch a public file or paste robots.txt for local analysis:
URL mode sends the host to Web Aloha's server. The endpoint may keep its response in instance memory for up to ten minutes.
Paste mode makes no API request. Analysis and path testing stay in this browser tab.
Found issues? That’s exactly what our SEO audit fixes.
Our SEO Audit + Essentials package ($555, one-time) finds and fixes exactly these kinds of problems, with a prioritized plan in plain English.
Looks clean here. How is the rest of your SEO?
One passing check doesn’t make a ranking strategy. We audit the full picture: content, technical SEO, internal links, and search intent.
Want us to act on this result? Fix these findings →
Prefer the partnership route? Refer clients to us, earn 10% →
How the Robots.txt Tester & Checker Works
Choose whether to fetch a public robots.txt file or inspect content you are still editing. The two modes use different data paths:
- Choose a source: URL mode makes the existing server request and checks host variants. Paste or Edit mode performs all analysis locally and makes no request.
- Review the file: inspect line-specific common syntax findings, User-agent groups, Allow and Disallow rules, and Sitemap references. Edit the text and rerun analysis before copying it.
- Test one path: enter a crawler product token or full user-agent string plus a URL or path. The browser applies the longest matching rule, supports
*and a final$, and lets Allow win an equal-specificity tie.
The simulator is a practical RFC-style diagnostic. Crawler-specific extensions, percent-encoding behavior, cached files, and implementation differences can produce another result in production.
What to Check in Your Robots.txt
A robots.txt file is easy to write and even easier to get wrong. When reviewing your results, focus on these essentials:
- Response and final host: an HTTP 200 response provides a file to evaluate. A 404 commonly means no special crawl rules. A 403, timeout, or server error needs separate verification because crawler handling varies by status and operator.
- Broad rules and exceptions:
Disallow: /underUser-agent: *is a broad block, but a longer Allow rule can create an exception. Test important paths rather than reading one line in isolation. - Sitemap directive: an absolute Sitemap URL can help participating crawlers discover a sitemap. It is optional and does not replace internal links or sitemap submission and monitoring.
- Critical resources: test page, CSS, JavaScript, image, and API paths required for rendering. Robots.txt controls crawl access, not whether a URL can appear in an index.
- Crawler-specific intent: training crawlers, search crawlers, and user-triggered fetchers can use separate product tokens. A rule for one token does not automatically control the others.
For a deeper look at how robots.txt connects to SEO, treat it as one crawl-control layer alongside status codes, canonical signals, meta robots, internal linking, sitemaps, and server logs.
A Practical Replacement for One Retired Testing Workflow
Google retired the old Search Console robots.txt tester. Its current reporting focuses on files Google has fetched and their status for properties you can access.
This page restores the useful edit-and-test workflow without claiming to reproduce Google's parser. You can fetch a public file, paste an unpublished draft, review common syntax findings, and test one user agent against one path.
For a high-impact change, verify the published response, crawler documentation, server logs, and relevant webmaster tooling. A local match is evidence about the text you tested, not a guarantee that every crawler has fetched the same file or applies every extension identically.
Use Separate Controls for Separate AI Crawlers
Robots.txt can express access preferences for named crawler tokens. It does not itself describe whether downstream systems train a model, build a search index, retrieve a page after a user request, rank an answer, or cite a source.
Operators publish different tokens for different purposes. Current examples include:
- GPTBot and OAI-SearchBot: OpenAI separates potential foundation-model training from automatic crawling used to surface pages in ChatGPT search. ChatGPT-User covers certain user-triggered requests, where OpenAI notes that robots.txt may not apply.
- ClaudeBot, Claude-SearchBot, and Claude-User: Anthropic separates potential model-development collection, search optimization, and user-directed retrieval.
- Google-Extended: a control token for specified Gemini training and grounding uses. It does not control inclusion or ranking in Google Search; regular Googlebot rules remain separate.
- Other operators: Common Crawl, Perplexity, and other services publish their own tokens and policies. Confirm the current documentation before deploying a rule.
Decide separately which uses fit your content, licensing, privacy, and visibility goals. Allowing a crawler does not guarantee retrieval, ranking, or citation, while blocking one token does not necessarily opt out of every product or user-triggered request.
After checking your robots.txt here, run your sitemap through the validator to catch format and availability issues. The AI Search Visibility Checker reviews additional public signals, while the llms.txt generator can produce an optional machine-readable content guide. Neither file guarantees discovery, retrieval, or citation.
Want the wider decision framework? Our article on robots.txt for AI search and GEO explains how to evaluate crawler tokens by use case. For deployment details, the guide on how to check and test robots.txt covers examples, wildcard patterns, and common mistakes.
Next steps
Robots.txt Tester & Validator related tools and articles
Continue with the closest follow-up checks and guides based on this tool's topic, crawl intent, and optimization workflow.
Robots.txt Tester, Validator & Checker: FAQ
Which robots.txt URL does the checker test?
What do the URL fetch summary badges validate?
How does the user-agent and path simulator decide a result?
What does Blocks crawling mean?
Is a missing robots.txt file an error?
Why can the report show cached: true?
What are the fetch limits and common failure modes?
What data is processed or retained?
Free 48-Hour Website Audit
Not sure what to fix first on your own website? We'll review it and tell you, in plain English. Free & non-obligatory.
Need Help With Your Robots.txt?
We audit crawl controls, diagnose conflicting directives, and connect robots.txt changes to technical SEO priorities.