Meta-ExternalAgent: Meta's AI training crawler and robots.txt
Meta-ExternalAgent crawls for uses such as AI model training and product indexing. Its user agent, robots.txt rule, 24-hour cache, and Meta's other agents.
- Operator
- Meta
- robots.txt token
meta-externalagent- User agents
meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)meta-externalagent/1.1- Purpose
- Model training
- Honours robots.txt
- Its operator documents robots.txt as the control
- Published IP ranges
- None published
- Reverse DNS
- None documented
- Official documentation
- Meta: Meta web crawlers
- Verified on
- 2026-10-05
What Meta-ExternalAgent does
Meta says Meta-ExternalAgent crawls the web for use cases such as training foundation AI models or improving products by indexing content directly. Meta says the user agent in server logs will be similar to one of the 2 strings above; its page prints the link as a relative path, so a log filter should match on meta-externalagent. [1]
Meta's robots.txt example writes the token in lowercase, meta-externalagent. RFC 9309 requires crawlers to match the token without regard to case; Meta does not say how its crawler matches, so writing it as Meta prints it is the safe choice. [1][2]
Allow or block Meta-ExternalAgent in robots.txt
Add one of these groups to the robots.txt at the root of each host (each scheme and port counts separately) it should apply to. A crawler that finds one or more groups naming its token follows those groups, merged into one, and ignores the User-agent: * group, so repeat in them any * rules you still want applied (RFC 9309, section 2.2.1).
User-agent: meta-externalagent
Allow: /User-agent: meta-externalagent
Disallow: /What blocking Meta-ExternalAgent changes
Meta's documented way to block its crawlers is a robots.txt Disallow naming the relevant crawler. Meta asks site owners to allow up to 24 hours for a change to take effect, because its crawlers may cache robots.txt for that long. [1]
Meta documents other agents with their own names, which a Meta-ExternalAgent rule does not address. Meta-WebIndexer helps Meta cite and link to content in Meta AI's responses. Meta-ExternalAds crawls for uses such as improving advertising and other business products. Meta-ExternalFetcher fetches individual links at a user's request and, Meta says, may bypass robots.txt rules. [1]
How to verify a request really is Meta-ExternalAgent
Meta's crawler page publishes no IP list, reverse-DNS hostname or other verification method for Meta-ExternalAgent. A request carrying its user agent can be matched by name, but the page gives no way to prove it came from Meta. [1]
A crawler policy is one part of being readable by assistants. The AI search readiness guide covers the rest, and robots.txt and crawl access covers the firewall and bot-protection rules that block crawlers before robots.txt is ever read. A full SEO report covers AI search readiness alongside the rest of a site's technical health.
Sources
- Meta: Meta web crawlers, verified on 2026-10-05
- IETF: RFC 9309, Robots Exclusion Protocol, verified on 2026-10-05
Get the complete diagnosis of your site
An evidence-backed report and a prioritized action plan, on a plan with monthly credits.