All AI crawlers

ClaudeBot: Anthropic's training crawler and its robots.txt rule

ClaudeBot collects content that may contribute to training Anthropic's models. How to allow or block it, the per-subdomain rule, Crawl-delay and the IP list.

Operator
Anthropic
robots.txt token
ClaudeBot
User agent
Not published by the operator
Purpose
Model training
Honours robots.txt
Yes, its operator states it does
Reverse DNS
None documented
Verified on
2026-10-05

What ClaudeBot does

Anthropic says ClaudeBot helps enhance the utility and safety of its generative AI models by collecting web content that could potentially contribute to their training. [1]

Anthropic names the robots.txt token but does not publish a full user-agent string, so rules and log filters should match on ClaudeBot. [1]

Allow or block ClaudeBot in robots.txt

Add one of these groups to the robots.txt at the root of each host (each scheme and port counts separately) it should apply to. A crawler that finds one or more groups naming its token follows those groups, merged into one, and ignores the User-agent: * group, so repeat in them any * rules you still want applied (RFC 9309, section 2.2.1).

Allow ClaudeBotrobots.txt
User-agent: ClaudeBot
Allow: /
Block ClaudeBotrobots.txt
User-agent: ClaudeBot
Disallow: /

What blocking ClaudeBot changes

Anthropic says restricting ClaudeBot signals that the site's future materials should be excluded from its AI model training datasets. [1]

Anthropic says its bots honor industry-standard robots.txt directives, support the non-standard Crawl-delay extension, and will not attempt to bypass CAPTCHAs. Each subdomain needs its own rule: Anthropic asks site owners to add the opt-out for every subdomain they want excluded. [1]

Anthropic lists Claude-SearchBot and Claude-User under their own tokens, so that owners can enable some and limit others. Under the robots.txt standard, a group that names only ClaudeBot does not apply to those 2. [1][3]

How to verify a request really is ClaudeBot

Anthropic publishes a list of IP addresses and says that a crawler whose source address is on it is coming from Anthropic. It also warns that blocking those addresses may not work correctly or persistently guarantee an opt-out, because doing so impedes Anthropic's ability to read robots.txt, and says opting out requires the robots.txt rule. [1][2]

A crawler policy is one part of being readable by assistants. The AI search readiness guide covers the rest, and robots.txt and crawl access covers the firewall and bot-protection rules that block crawlers before robots.txt is ever read. A full SEO report covers AI search readiness alongside the rest of a site's technical health.

Sources

  1. Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?, verified on 2026-10-05
  2. Anthropic: published crawler IP addresses (JSON), verified on 2026-10-05
  3. IETF: RFC 9309, Robots Exclusion Protocol, verified on 2026-10-05

Get the complete diagnosis of your site

An evidence-backed report and a prioritized action plan, on a plan with monthly credits.