ClaudeBot: Anthropic's training crawler and its robots.txt rule
ClaudeBot collects content that may contribute to training Anthropic's models. How to allow or block it, the per-subdomain rule, Crawl-delay and the IP list.
- Operator
- Anthropic
- robots.txt token
ClaudeBot- User agent
- Not published by the operator
- Purpose
- Model training
- Honours robots.txt
- Yes, its operator states it does
- Published IP ranges
- https://claude.com/crawling/bots.json
- Reverse DNS
- None documented
- Official documentation
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Verified on
- 2026-10-05
What ClaudeBot does
Anthropic says ClaudeBot helps enhance the utility and safety of its generative AI models by collecting web content that could potentially contribute to their training. [1]
Anthropic names the robots.txt token but does not publish a full user-agent string, so rules and log filters should match on ClaudeBot. [1]
Allow or block ClaudeBot in robots.txt
Add one of these groups to the robots.txt at the root of each host (each scheme and port counts separately) it should apply to. A crawler that finds one or more groups naming its token follows those groups, merged into one, and ignores the User-agent: * group, so repeat in them any * rules you still want applied (RFC 9309, section 2.2.1).
User-agent: ClaudeBot
Allow: /User-agent: ClaudeBot
Disallow: /What blocking ClaudeBot changes
Anthropic says restricting ClaudeBot signals that the site's future materials should be excluded from its AI model training datasets. [1]
Anthropic says its bots honor industry-standard robots.txt directives, support the non-standard Crawl-delay extension, and will not attempt to bypass CAPTCHAs. Each subdomain needs its own rule: Anthropic asks site owners to add the opt-out for every subdomain they want excluded. [1]
Anthropic lists Claude-SearchBot and Claude-User under their own tokens, so that owners can enable some and limit others. Under the robots.txt standard, a group that names only ClaudeBot does not apply to those 2. [1][3]
How to verify a request really is ClaudeBot
Anthropic publishes a list of IP addresses and says that a crawler whose source address is on it is coming from Anthropic. It also warns that blocking those addresses may not work correctly or persistently guarantee an opt-out, because doing so impedes Anthropic's ability to read robots.txt, and says opting out requires the robots.txt rule. [1][2]
A crawler policy is one part of being readable by assistants. The AI search readiness guide covers the rest, and robots.txt and crawl access covers the firewall and bot-protection rules that block crawlers before robots.txt is ever read. A full SEO report covers AI search readiness alongside the rest of a site's technical health.
Sources
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?, verified on 2026-10-05
- Anthropic: published crawler IP addresses (JSON), verified on 2026-10-05
- IETF: RFC 9309, Robots Exclusion Protocol, verified on 2026-10-05
Get the complete diagnosis of your site
An evidence-backed report and a prioritized action plan, on a plan with monthly credits.