Google-Extended
Google-Extended is an AI data scraper run by Google.
- Operator
- Type
- AI Data Scraper
- Official page
- developers.google.com
- robots.txt token
Google-Extended- Follows robots.txt
- Yes, according to ai.robots.txt.
- In Vireo
- Excluded, reported under AI crawlers
What it does
A robots.txt control token, not a separate crawler. It decides whether content Google crawls can be used to train Gemini models and for grounding in Gemini Apps and Vertex AI.
Google-Extended is a standalone product token that web publishers can use to manage whether content Google crawls from their sites may be used for training future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini and for grounding (providing content from the Google Search index to the model at prompt time to improve factuality and relevancy) in Gemini Apps and Grounding with Google Search on Vertex AI.
Google
User agent
Requests from Google-Extended carry a user agent like this:
Google-ExtendedDevice Detector identifies it with this pattern (a case-insensitive regular expression):
Google-ExtendedAnyone can put any name in a user agent. Where the operator publishes IP ranges, checking the request's address against them is the way to confirm a visit is genuine.
How to block or allow Google-Extended
Add a group for its token to the robots.txt file at the root of your site. To keep it out of the whole site:
User-agent: Google-Extended
Disallow: /To let it crawl everything, use an empty Disallow: or Allow: / instead. WordPress serves a virtual robots.txt that many SEO plugins let you edit. A real robots.txt file uploaded to the site root takes its place.
User-agent token in robots.txt Google-Extended
Google
robots.txt is a request, so it only works on crawlers that honour it. A bot that ignores it has to be blocked at the server or firewall, by user agent or IP address.
Worth knowing
It has no user agent string of its own, so it never shows up in logs. Blocking it doesn't affect Google Search inclusion or ranking.
Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity.
Google
How Vireo counts Google-Extended
Vireo keeps requests from Google-Extended out of your pageviews and visitors and counts them on its bot report under AI crawlers, so you can see how much it left out rather than wonder why your server logs show more traffic.
Because it's an AI crawler, its requests also count toward the pages AI crawlers took, which Vireo shows next to the readers AI assistants sent back. See how the AI assistants card works.
Checked by running Vireo Analytics' own bot filter over Google-Extended's user agents.
Sources
- Google: crawler documentation (checked October 2026)
- Matomo Device Detector, bots.yml (LGPL-3.0)
- ai.robots.txt (MIT)
Other bots from Google
More ai crawlers
- Ai2Bot (Allen Institute for AI (Ai2))
- Ai2Bot-DeepResearchEval (The Allen Institute for Artificial Intelligence)
- Ai2Bot-Dolma (The Allen Institute for Artificial Intelligence)
- Amazonbot (Amazon)
- Amazonbot-Video (Amazon.com, Inc.)
- AmazonBuyForMe (Amazon.com, Inc.)
- Amzn-SearchBot (Amazon.com, Inc.)
- Amzn-User (Amazon.com, Inc.)
- Andibot (Andi)
- Anthropic AI (Anthropic, PBC)
See the bots on your own site
Vireo Analytics is a free WordPress plugin. It keeps over 800 known crawlers out of your numbers and shows you what it filtered, by category.
Get Vireo free on WordPress.org