AI companies run several crawlers, and they do different jobs. Some collect data to train models. Others fetch pages live to answer a user question, or power AI search. If you block everything you may disappear from AI answers, and if you allow everything you give up control. This guide lists the main crawlers and how to choose.
Three kinds of AI crawler
- Training crawlers collect content to train models.
- Search crawlers index pages so an AI search product can cite them.
- User-triggered fetchers load a page because a person asked an assistant about it.
Common user agents
| Company | Agent | Typical purpose |
|---|---|---|
| OpenAI | GPTBot | Training data |
| OpenAI | OAI-SearchBot | Indexing for ChatGPT search |
| OpenAI | ChatGPT-User | User-triggered fetches |
| Anthropic | ClaudeBot | Training data |
| Anthropic | Claude-SearchBot, Claude-User | Search indexing and user-triggered fetches |
| Perplexity | PerplexityBot, Perplexity-User | Indexing and user-triggered fetches |
Google-Extended | Control token for Gemini use of content | |
| Apple | Applebot-Extended | Control token for Apple AI training |
Providers change names and behaviour. Check each company's current documentation before you edit your rules.
How to choose
Ask what you want from AI. If you want to be cited in AI answers, allow the search and user-triggered crawlers. Whether to allow training crawlers is a business decision: it can help models know your brand, but you give your content away. Many publishers allow search bots and block training bots. Software and e-commerce brands usually benefit from being known, so I tend to allow all of them on public marketing pages.
Example robots.txt
User-agent: * Allow: / # Allow AI search and user-triggered fetches User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: PerplexityBot Allow: / # Example: opt out of training only User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / Sitemap: https://example.com/sitemap.xml
Test your rules
- Open
/robots.txtand confirm it loads with status 200. - Check there is no accidental
Disallow: /for all agents. - Make sure a firewall or CDN is not blocking these bots separately.
- Confirm important pages are server-rendered so crawlers see the text.
Frequently asked questions
Will blocking GPTBot remove me from ChatGPT answers?
Blocking GPTBot stops OpenAI using your pages for model training. Live answers in ChatGPT search rely on OAI-SearchBot and user-initiated fetches by ChatGPT-User, which are separate.
Does Google-Extended affect my Google rankings?
No. Google-Extended is a control token for use of your content by Gemini and related AI features. It does not affect Search ranking.
Can robots.txt fully hide content from AI?
No. It is a request that well-behaved crawlers respect, not a security control. Use authentication for private content.
Written by Hiral Doshi, digital marketer in Mumbai. Last updated 2026-10-02.