Googlebot

Googlebot is a search engine crawler run by Google.

Operator
Google
Type
Search bot
Official page
developers.google.com
robots.txt token
Googlebot
Follows robots.txt
Yes, according to Google.
Runs JavaScript
Yes, according to Google.
Published IP ranges
developers.google.com/static/crawling/ipranges/common-crawlers.json
In Vireo
Excluded, reported under Search engines

What it does

Google's main crawler for Google Search, with a smartphone and a desktop variant.

Googlebot is the generic name for two types of web crawlers used by Google Search

Google

User agent

Requests from Googlebot carry user agents like these:

Googlebot-Image/1.0
Googlebot-Video/1.0
Googlebot/Nutch-1.7

Device Detector identifies it with these patterns (a case-insensitive regular expression):

Adwords-(?:DisplayAds|Express|Instant)|Google Web Preview|Google[ -]Publisher[ -]Plugin|Google-(?:adstxt|Ads-Conversions|Ads-Qualify|Adwords|AMPHTML|Assess|BusinessLinkVerification|HotelAdsVerifier|InspectionTool|Lens|PageRenderer|Shopping-Quality|Sites-Thumbnails|speakr|Stale-Content-Probe|Test|Youtube-Links)|(?:AdsBot|APIs|Feedfetcher|Mediapartners)-Google(?:-Mobile)?|Google(?:AdSenseInfeed|AssociationService|bot|Other|Prober|Producer|Sites)|Google.*/\+/web/snippet
^Google$

Anyone can put any name in a user agent. Where the operator publishes IP ranges, checking the request's address against them is the way to confirm a visit is genuine.

How to block or allow Googlebot

Add a group for its token to the robots.txt file at the root of your site. To keep it out of the whole site:

User-agent: Googlebot
Disallow: /

To let it crawl everything, use an empty Disallow: or Allow: / instead. WordPress serves a virtual robots.txt that many SEO plugins let you edit. A real robots.txt file uploaded to the site root takes its place.

both crawler types obey the same product token (user agent token) in robots.txt, and so you cannot selectively target either Googlebot Smartphone or Googlebot Desktop using robots.txt.

Google

They always obey robots.txt rules when crawling automatically.

Google

robots.txt is a request, so it only works on crawlers that honour it. A bot that ignores it has to be blocked at the server or firewall, by user agent or IP address.

Worth knowing

Blocking it in robots.txt stops crawling but doesn't keep a URL out of results. Google says to use noindex for that.

blocking Googlebot from crawling a page doesn't prevent the URL of the page from appearing in search results

Google

Does Googlebot run JavaScript?

Yes. Google says it does.

Once Google's resources allow, a headless Chromium renders the page and executes the JavaScript.

Google

How Vireo counts Googlebot

Vireo keeps requests from Googlebot out of your pageviews and visitors and counts them on its bot report under Search engines, so you can see how much it left out rather than wonder why your server logs show more traffic.

Checked by running Vireo Analytics' own bot filter over Googlebot's user agents.

Sources

Other bots from Google

More search engines

All search engines

See the bots on your own site

Vireo Analytics is a free WordPress plugin. It keeps over 800 known crawlers out of your numbers and shows you what it filtered, by category.

Get Vireo free on WordPress.org