Contact

WritingGuide

Robots.txt for AI bots: who to allow

By Rues · Published · Updated

Each AI company now runs more than one crawler, and each crawler has a different job. One builds a search index, one collects training data, and one fetches a page when a user asks about it. You can allow or block each of them separately in robots.txt. This guide lists the bots that matter today, what their owners say each one does, and the robots.txt groups we would write for the common choices.

Which AI bots exist, and what does each one do?

The table below is built only from each company's own documentation, read on 7 October 2026. Where a document is silent on something, we leave the cell as "not stated".

CompanyRobots.txt tokenJobFollows robots.txt?
OpenAIOAI-SearchBotSurfaces sites in ChatGPT search resultsYes; OpenAI says changes take about 24 hours
OpenAIGPTBotCrawls content that may be used to train foundation modelsYes; disallowing it opts out of training
OpenAIChatGPT-UserFetches pages for user actions in ChatGPT"Robots.txt rules may not apply"
AnthropicClaude-SearchBotIndexes content to improve search resultsYes
AnthropicClaudeBotCollects content that may be used for trainingYes
AnthropicClaude-UserFetches pages when a user asks Claude a questionYes
PerplexityPerplexityBotSurfaces and links sites in Perplexity search; not used for trainingYes; Perplexity says webmasters manage it with robots.txt tags and recommends allowing it
PerplexityPerplexity-UserVisits a page to answer a user's question"Generally ignores robots.txt rules"
GoogleGooglebotMain crawler for Google SearchYes
GoogleGoogle-ExtendedControls use of content for Gemini training and groundingToken only, no separate crawler

OpenAI states that its settings are independent of each other: you can allow OAI-SearchBot and block GPTBot. Anthropic also describes three agents with separate jobs, each with its own name in robots.txt. Blocking Claude-SearchBot stops your content from being indexed for search, and blocking Claude-User stops Claude from fetching your page when a user asks about it.

What is the difference between a search bot and a training bot?

A search bot decides whether your page can appear, with a link, inside an AI answer. A training bot collects pages that may end up in the data a future model learns from. The first affects whether you can be cited today. The second is a choice about how your content is used.

Google-Extended needs a note of its own. Google says it has no separate user agent; Google crawls with its usual crawlers, and the token in robots.txt only tells Google whether that content may be used for Gemini training and for grounding in Gemini Apps and Vertex AI. Google also says Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal". The page does not mention AI Overviews or AI Mode, so we make no claim about them here.

What about user-triggered fetchers?

ChatGPT-User, Claude-User and Perplexity-User come to your site because a person asked a question. OpenAI writes that robots.txt rules "may not apply" to ChatGPT-User. Perplexity writes that Perplexity-User "generally ignores robots.txt rules". Anthropic says its bots respect robots.txt, including Claude-User.

In practice, a robots.txt block will not reliably keep every user-triggered fetch away. If a page must stay private, put it behind a login.

The trap: a specific group replaces the * group

A crawler follows one group of rules. Google's documentation says crawlers pick the group with the most specific matching user agent, and "user agent specific groups and global groups (*) are not combined". RFC 9309, the robots.txt standard, says the same: the * group applies only when no group matches the crawler.

So this file opens your admin area to GPTBot:

User-agent: *
Disallow: /panel/

User-agent: GPTBot
Allow: /

GPTBot now reads only its own group, and Disallow: /panel/ is gone for it. Repeat the shared rules in every named group, or list several user agents above one set of rules.

Robots.txt examples for three common choices

1. Open to search and training. This is our own choice for ruesandora.com, made on 7 October 2026: all crawlers allowed, training bots included. Our file has no private area, so it is only User-agent: *, Allow: / and the sitemap line. With an admin area it looks like this:

User-agent: *
Disallow: /panel/

Sitemap: https://example.com/sitemap.xml

2. Open to search, closed to training. Search bots and Googlebot fall through to the * group. The training bots get their own group.

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /

User-agent: *
Disallow: /panel/

Sitemap: https://example.com/sitemap.xml

3. Closed to AI search and training, open to Google Search.

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Google-Extended
Disallow: /

User-agent: *
Disallow: /panel/

The third file will not stop ChatGPT-User or Perplexity-User for the reasons above. It also removes you from the AI search results of those companies, so you will not be cited there.

Changes are not instant. OpenAI says about 24 hours for search, Perplexity says up to 24 hours, and RFC 9309 says crawlers should not use a cached robots.txt for more than 24 hours.

Why your robots.txt may not be the whole story

Robots.txt is a request. Cloudflare's own documentation puts it plainly: robots.txt "does not prevent crawlers from accessing your content at a technical level." The reverse is also true. A CDN can block a bot that your robots.txt allows, and the bot never reaches your file.

On Cloudflare there are three settings to look at:

  • AI Crawl Control. You can allow or block each AI crawler. A block creates a WAF custom rule on your zone, and the blocked crawler gets a 403 Forbidden or 402 Payment Required response.
  • Managed robots.txt. When turned on, Cloudflare puts its own rules in front of your robots.txt and serves both as one file. Those rules disallow GPTBot, ClaudeBot, Google-Extended and several others. What crawlers read can differ from the file in your repository.
  • AI bot policies and new defaults. Cloudflare retired its older "Block AI bots" setting on 15 September 2026. Since that date, new domains get defaults that block bots classified as Training or Agent on pages that display ads, while Search stays allowed.

A request sent with a fake user agent from your own machine does not show what the real crawler, coming from its published IP ranges, receives. OpenAI, Anthropic and Perplexity publish IP lists for their bots.

A short checklist

  1. Write down your decision first: search yes or no, training yes or no.
  2. Write robots.txt groups that match the decision, and repeat shared Disallow lines in every named group.
  3. Open https://yoursite.com/robots.txt in a browser and read what is actually served.
  4. Check your CDN's bot settings, including any managed robots.txt and default policies.
  5. After a few days, look at your server or CDN logs for the bots you allowed.

If you want a second pair of eyes on your bot settings, write to us.

Sources

Want to talk about this for your brand?

Write about your project and what you want to do. The message goes straight to Rues.

Open the contact form