Robots.txt for AI Search: Crawler Controls Without the Hype

Author: Lucky Oleg | Published Updated
Robots.txt for AI Search: Crawler Controls Without the Hype

Robots.txt decisions should be based on named crawlers and documented product effects. “Allow all AI bots” and “block all AI bots” are both too blunt for most sites.

The key is to separate three activities:

  • automated crawling for search or retrieval indexes;
  • collection for potential model training;
  • a fetch triggered by a user request.

The same company can use different tokens for each purpose.

Platform controls at a glance

PlatformTokenDocumented purposeImportant limitation
GoogleGooglebotGoogle Search crawling, including AI Overviews and AI ModeAllowing crawling does not guarantee indexing or display
GoogleGoogle-ExtendedControl over Gemini model training and grounding in certain other Google systemsDoes not affect Google Search inclusion or ranking
OpenAIOAI-SearchBotDiscovery for ChatGPT search summaries, snippets, citations, and linksAccess is necessary for full content use, not a placement guarantee
OpenAIGPTBotPotential model-training collectionSeparate from OAI-SearchBot
PerplexityPerplexityBotCrawling to surface and link sites in Perplexity search resultsAccess does not guarantee a result
PerplexityPerplexity-UserUser-triggered page visitsPerplexity says it generally ignores robots.txt

Sources: Google AI features, Google common crawlers, OpenAI publisher FAQ, and Perplexity crawler documentation.

Crawler documentation can change. Recheck official pages before making a policy decision.

Google Search and Google-Extended are different controls

For AI Overviews and AI Mode in Google Search, Google says the relevant crawler control is Googlebot. Pages also need to be indexed and eligible to appear with a snippet. There is no separate AI Overview crawler to allow.

Google-Extended is a standalone robots token used to control whether Google-crawled content may support future Gemini model training and grounding in specified Gemini and Vertex AI products. Google explicitly says this token does not affect inclusion or ranking in Google Search.

Therefore, this is a valid policy:

User-agent: Googlebot
Allow: /

User-agent: Google-Extended
Disallow: /

It keeps Google Search crawl access while declining the documented Google-Extended uses. Whether that is the right choice is a business and publishing decision, not an SEO universal.

OpenAI search and training are separate

OpenAI tells publishers who want content included in ChatGPT summaries and snippets to allow OAI-SearchBot. It uses GPTBot as the separate potential-training control.

For example, a publisher could allow ChatGPT search while opting out of potential training:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

OpenAI also notes that a disallowed page’s link and title may sometimes be surfaced when the URL is learned elsewhere. If the intent is to prevent indexing rather than only crawling, use the platform’s documented indexing controls and remember that a crawler must access a page to see a page-level noindex directive.

Perplexity has an automated crawler and a user fetcher

Perplexity recommends allowing PerplexityBot for sites that want to appear in its search results. It documents Perplexity-User separately as a user-triggered fetcher and says that fetcher generally ignores robots.txt.

This is why a robots audit should not report every observed user agent as an equivalent opt-in control. Some are automated crawlers; others represent a user’s request.

Do not copy a universal allow list

Before changing rules, answer these questions:

  1. Is the content public, licensed, paywalled, private, or regulated?
  2. Is the goal search discovery, direct referrals, model-training control, or server-load control?
  3. Which crawler does the platform officially associate with that goal?
  4. Does a CDN or bot-management layer block verified crawlers even when robots.txt allows them?
  5. Are there sensitive URLs that should require authentication rather than rely on robots.txt?

Robots.txt is public and is not a security boundary. Never use it to protect confidential data. Use authentication and authorization.

A conservative public-site example

This example keeps public search crawlers open and blocks an internal area. It intentionally does not prescribe a training policy:

User-agent: Googlebot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: *
Disallow: /internal/

Sitemap: https://example.com/sitemap.xml

Only include /internal/ if that route exists, and still protect it with authentication if it contains anything private.

How to audit safely

  1. Fetch the live /robots.txt, not only the repository copy.
  2. Check status code, redirects, content type, and caching.
  3. Evaluate each important URL against each named crawler group.
  4. Inspect CDN, WAF, and server logs for verified crawler failures.
  5. Verify crawler identity using official IP or reverse-DNS guidance where available; user-agent strings can be spoofed.
  6. Test indexing and preview controls separately from crawl permission.
  7. Record the policy decision and date so future maintainers know why each rule exists.

Use the robots.txt tester for parsing and the general robots.txt guide for syntax. A passing test means the file can be interpreted; it does not prove citation, indexing, or crawler identity.

How this fits into GEO

Crawler access is only an eligibility condition. After access, platforms still decide whether to index, retrieve, trust, and cite a page. Helpful content, internal discovery, accurate schema, platform policies, and query relevance remain separate factors.

For the wider workflow, see the GEO guide and AI visibility measurement guide. Make access choices deliberately, measure what happens, and avoid turning crawler permission into a result promise.

Recommended tools

Recommended tools for this guide

Use these free tools to apply the ideas from this guide to your own website.

Useful info? Spread the Aloha:

Lucky Oleg

Lucky Oleg is the founder of Web Aloha, a web design & SEO agency helping businesses ride the digital wave. With years of experience in WordPress, technical SEO, and web performance, he writes about what actually works in the real world.