All AI crawlers

GPTBot: OpenAI's training crawler, user agent and robots.txt

What GPTBot collects, its published user agent, how to allow or block it in robots.txt, and why blocking it does not opt a site out of ChatGPT search.

Operator
OpenAI
robots.txt token
GPTBot
User agent
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
Purpose
Model training
Honours robots.txt
Its operator documents robots.txt as the control
Reverse DNS
None documented
Verified on
2026-10-05

What GPTBot does

OpenAI uses GPTBot to crawl content that may be used in training its generative AI foundation models, and describes it as helping make those models more useful and safe. It reads robots.txt under the token GPTBot. [1]

OpenAI publishes the user agent as an example whose version number may change, and says that when GPTBot fetches robots.txt it may add a robots.txt marker to the string. A log filter should match the GPTBot token rather than the full string. [1]

Allow or block GPTBot in robots.txt

Add one of these groups to the robots.txt at the root of each host (each scheme and port counts separately) it should apply to. A crawler that finds one or more groups naming its token follows those groups, merged into one, and ignores the User-agent: * group, so repeat in them any * rules you still want applied (RFC 9309, section 2.2.1).

Allow GPTBotrobots.txt
User-agent: GPTBot
Allow: /
Block GPTBotrobots.txt
User-agent: GPTBot
Disallow: /

What blocking GPTBot changes

OpenAI states that disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models. [1]

ChatGPT search has its own token. OpenAI states that each of its robots.txt settings is independent of the others, so a site can disallow GPTBot and still allow OAI-SearchBot, the crawler that surfaces sites in ChatGPT search. [1]

Allowing both need not double the crawling: when a site allows both bots, OpenAI says it may use the results from 1 crawl for both uses. [1]

How to verify a request really is GPTBot

OpenAI publishes the IP addresses GPTBot uses as a JSON list. A request that calls itself GPTBot from an address outside that list does not come from a range OpenAI publishes for it. OpenAI documents no reverse-DNS hostname, so the list is the check. [1][2]

A crawler policy is one part of being readable by assistants. The AI search readiness guide covers the rest, and robots.txt and crawl access covers the firewall and bot-protection rules that block crawlers before robots.txt is ever read. A full SEO report covers AI search readiness alongside the rest of a site's technical health.

Sources

  1. OpenAI: Overview of OpenAI crawlers, verified on 2026-10-05
  2. OpenAI: GPTBot IP addresses (JSON), verified on 2026-10-05

Get the complete diagnosis of your site

An evidence-backed report and a prioritized action plan, on a plan with monthly credits.