Google-Extended

Google-Extended is an AI data scraper run by Google.

Operator
Google
Type
AI Data Scraper
Official page
developers.google.com
robots.txt token
Google-Extended
Follows robots.txt
Yes, according to ai.robots.txt.
In Vireo
Excluded, reported under AI crawlers

What it does

A robots.txt control token, not a separate crawler. It decides whether content Google crawls can be used to train Gemini models and for grounding in Gemini Apps and Vertex AI.

Google-Extended is a standalone product token that web publishers can use to manage whether content Google crawls from their sites may be used for training future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini and for grounding (providing content from the Google Search index to the model at prompt time to improve factuality and relevancy) in Gemini Apps and Grounding with Google Search on Vertex AI.

Google

User agent

Requests from Google-Extended carry a user agent like this:

Google-Extended

Device Detector identifies it with this pattern (a case-insensitive regular expression):

Google-Extended

Anyone can put any name in a user agent. Where the operator publishes IP ranges, checking the request's address against them is the way to confirm a visit is genuine.

How to block or allow Google-Extended

Add a group for its token to the robots.txt file at the root of your site. To keep it out of the whole site:

User-agent: Google-Extended
Disallow: /

To let it crawl everything, use an empty Disallow: or Allow: / instead. WordPress serves a virtual robots.txt that many SEO plugins let you edit. A real robots.txt file uploaded to the site root takes its place.

User-agent token in robots.txt Google-Extended

Google

robots.txt is a request, so it only works on crawlers that honour it. A bot that ignores it has to be blocked at the server or firewall, by user agent or IP address.

Worth knowing

It has no user agent string of its own, so it never shows up in logs. Blocking it doesn't affect Google Search inclusion or ranking.

Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity.

Google

How Vireo counts Google-Extended

Vireo keeps requests from Google-Extended out of your pageviews and visitors and counts them on its bot report under AI crawlers, so you can see how much it left out rather than wonder why your server logs show more traffic.

Because it's an AI crawler, its requests also count toward the pages AI crawlers took, which Vireo shows next to the readers AI assistants sent back. See how the AI assistants card works.

Checked by running Vireo Analytics' own bot filter over Google-Extended's user agents.

Sources

Other bots from Google

More ai crawlers

All ai crawlers

See the bots on your own site

Vireo Analytics is a free WordPress plugin. It keeps over 800 known crawlers out of your numbers and shows you what it filtered, by category.

Get Vireo free on WordPress.org