All AI crawlers

Meta-ExternalAgent: Meta's AI training crawler and robots.txt

Meta-ExternalAgent crawls for uses such as AI model training and product indexing. Its user agent, robots.txt rule, 24-hour cache, and Meta's other agents.

Operator
Meta
robots.txt token
meta-externalagent
User agents
meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)meta-externalagent/1.1
Purpose
Model training
Honours robots.txt
Its operator documents robots.txt as the control
Published IP ranges
None published
Reverse DNS
None documented
Official documentation
Meta: Meta web crawlers
Verified on
2026-10-05

What Meta-ExternalAgent does

Meta says Meta-ExternalAgent crawls the web for use cases such as training foundation AI models or improving products by indexing content directly. Meta says the user agent in server logs will be similar to one of the 2 strings above; its page prints the link as a relative path, so a log filter should match on meta-externalagent. [1]

Meta's robots.txt example writes the token in lowercase, meta-externalagent. RFC 9309 requires crawlers to match the token without regard to case; Meta does not say how its crawler matches, so writing it as Meta prints it is the safe choice. [1][2]

Allow or block Meta-ExternalAgent in robots.txt

Add one of these groups to the robots.txt at the root of each host (each scheme and port counts separately) it should apply to. A crawler that finds one or more groups naming its token follows those groups, merged into one, and ignores the User-agent: * group, so repeat in them any * rules you still want applied (RFC 9309, section 2.2.1).

Allow Meta-ExternalAgentrobots.txt
User-agent: meta-externalagent
Allow: /
Block Meta-ExternalAgentrobots.txt
User-agent: meta-externalagent
Disallow: /

What blocking Meta-ExternalAgent changes

Meta's documented way to block its crawlers is a robots.txt Disallow naming the relevant crawler. Meta asks site owners to allow up to 24 hours for a change to take effect, because its crawlers may cache robots.txt for that long. [1]

Meta documents other agents with their own names, which a Meta-ExternalAgent rule does not address. Meta-WebIndexer helps Meta cite and link to content in Meta AI's responses. Meta-ExternalAds crawls for uses such as improving advertising and other business products. Meta-ExternalFetcher fetches individual links at a user's request and, Meta says, may bypass robots.txt rules. [1]

How to verify a request really is Meta-ExternalAgent

Meta's crawler page publishes no IP list, reverse-DNS hostname or other verification method for Meta-ExternalAgent. A request carrying its user agent can be matched by name, but the page gives no way to prove it came from Meta. [1]

A crawler policy is one part of being readable by assistants. The AI search readiness guide covers the rest, and robots.txt and crawl access covers the firewall and bot-protection rules that block crawlers before robots.txt is ever read. A full SEO report covers AI search readiness alongside the rest of a site's technical health.

Sources

  1. Meta: Meta web crawlers, verified on 2026-10-05
  2. IETF: RFC 9309, Robots Exclusion Protocol, verified on 2026-10-05

Get the complete diagnosis of your site

An evidence-backed report and a prioritized action plan, on a plan with monthly credits.