Reference
AI crawlers and their robots.txt tokens
The main AI crawlers' and control tokens' robots.txt names, user agents and purposes, from each operator's own documentation, re-verified monthly.
Each operator below names its agents with their own robots.txt tokens, though some say their user-triggered fetchers may not follow robots.txt. Most operators listed here run more than one agent, typically some mix of a crawler that puts sites into answers, an agent that fetches a page when someone asks, and a crawler or token for model training. Each page covers one of them, cites the operator's own documentation, and gives the robots.txt lines to allow or block it.
How often real sites block them is a separate question, with its own sample and method: what the robots.txt files of 45 audited sites allowed.
Crawlers that put sites into AI search answers
OpenAI
OAI-SearchBot
OpenAI says sites that disallow OAI-SearchBot are not shown in ChatGPT search answers. Its user agents, robots.txt rule, IP list and the roughly 24-hour lag.
Anthropic
Claude-SearchBot
Claude-SearchBot is the Anthropic crawler that, it says, improves search result quality for users. What a Disallow costs, and how it differs from ClaudeBot.
Perplexity
PerplexityBot
Perplexity says PerplexityBot links sites in its search results and is not used to crawl content for AI foundation models. Its robots.txt rule and IP list.
Agents that fetch a page when someone asks
OpenAI
ChatGPT-User
ChatGPT-User may visit a page when a person asks ChatGPT or a custom GPT a question. Why OpenAI says robots.txt may not apply, and which token controls search.
Anthropic
Claude-User
Claude-User may fetch a page when someone asks Claude a question. Anthropic says its bots honor robots.txt; OpenAI and Perplexity say otherwise for theirs.
Perplexity
Perplexity-User
Perplexity-User may visit a page to answer a person's question and link to it. The 2 things Perplexity says about robots.txt for it, and how to verify it.
Crawlers that collect model-training data
OpenAI
GPTBot
What GPTBot collects, its published user agent, how to allow or block it in robots.txt, and why blocking it does not opt a site out of ChatGPT search.
Anthropic
ClaudeBot
ClaudeBot collects content that may contribute to training Anthropic's models. How to allow or block it, the per-subdomain rule, Crawl-delay and the IP list.
Meta
Meta-ExternalAgent
Meta-ExternalAgent crawls for uses such as AI model training and product indexing. Its user agent, robots.txt rule, 24-hour cache, and Meta's other agents.
Amazon
Amazonbot
Amazonbot improves Amazon's services and may train its AI models. Its robots.txt rule, the 2 timings Amazon gives for a change, and the noarchive opt-out.
Tokens that govern how an existing crawler's data is used
Google-Extended
Google-Extended manages whether Google may use a site's content to train and ground Gemini. No user agent of its own; Google says Search is unaffected.
Apple
Applebot-Extended
Applebot-Extended lets a site keep its content out of Apple's foundation-model training. With Applebot allowed, the site stays in Siri, Spotlight and Safari.
Open web archives
Get the complete diagnosis of your site
An evidence-backed report and a prioritized action plan, on a plan with monthly credits.