Reference

AI crawlers and their robots.txt tokens

The main AI crawlers' and control tokens' robots.txt names, user agents and purposes, from each operator's own documentation, re-verified monthly.

Each operator below names its agents with their own robots.txt tokens, though some say their user-triggered fetchers may not follow robots.txt. Most operators listed here run more than one agent, typically some mix of a crawler that puts sites into answers, an agent that fetches a page when someone asks, and a crawler or token for model training. Each page covers one of them, cites the operator's own documentation, and gives the robots.txt lines to allow or block it.

How often real sites block them is a separate question, with its own sample and method: what the robots.txt files of 45 audited sites allowed.

Crawlers that put sites into AI search answers

Agents that fetch a page when someone asks

Crawlers that collect model-training data

Tokens that govern how an existing crawler's data is used

Open web archives

Get the complete diagnosis of your site

An evidence-backed report and a prioritized action plan, on a plan with monthly credits.