Google-Extended: the robots.txt token for Gemini training
Google-Extended manages whether Google may use a site's content to train and ground Gemini. No user agent of its own; Google says Search is unaffected.
- Operator
- robots.txt token
Google-Extended- User agent
- No request user agent of its own; fetching is done by Google's existing crawlers
- Purpose
- Usage control token (no crawler of its own)
- Honours robots.txt
- Its operator documents robots.txt as the control
- Published IP ranges
- https://developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests
- Reverse DNS
googlebot.com, google.com or googleusercontent.com (Google's check for any of its crawlers)- Official documentation
- Google: Google's common crawlers
- Verified on
- 2026-10-05
What Google-Extended does
Google describes Google-Extended as a standalone product token that lets a site manage whether content Google crawls from it may be used for training future generations of the Gemini models that power Gemini Apps and the Vertex AI API for Gemini. [1]
Google says the same token manages whether that content may be used for grounding in Gemini Apps and in Grounding with Google Search on Vertex AI: providing content from the Google Search index to the model at prompt time. [1]
Google-Extended is not a separate crawler. Google states that it has no separate HTTP user agent string; crawling is done with Google's existing user agents, and the token is used in a control capacity. [1]
Allow or block Google-Extended in robots.txt
Add one of these groups to the robots.txt at the root of each host (each scheme and port counts separately) it should apply to. This group controls how Google may use content its crawler collects. It does not change what Google's existing crawlers may fetch; that is decided by whichever group applies to the crawler doing the fetching.
User-agent: Google-Extended
Allow: /User-agent: Google-Extended
Disallow: /What blocking Google-Extended changes
Google states that Google-Extended does not affect a site's inclusion in Google Search and is not used as a ranking signal in Google Search. [1]
How to verify a request really is Google-Extended
Google does not name which of its crawlers fetch the content this token governs. Its standard check for any of its crawlers is a reverse DNS lookup ending in googlebot.com, google.com or googleusercontent.com, confirmed by a forward lookup back to the same address, or a match against its published IP lists. Its common crawlers resolve under googlebot.com. [1][2][3]
A crawler policy is one part of being readable by assistants. The AI search readiness guide covers the rest, and robots.txt and crawl access covers the firewall and bot-protection rules that block crawlers before robots.txt is ever read. A full SEO report covers AI search readiness alongside the rest of a site's technical health.
Sources
- Google: Google's common crawlers, verified on 2026-10-05
- Google: Verify requests from Google's crawlers and fetchers, verified on 2026-10-05
- Google: common crawler IP ranges (JSON), verified on 2026-10-05
Get the complete diagnosis of your site
An evidence-backed report and a prioritized action plan, on a plan with monthly credits.