YandexBot

YandexBot is a search engine crawler run by Yandex.

Operator
Yandex
Type
Search bot
Official page
yandex.com
robots.txt token
YandexBot
Follows robots.txt
Yes, according to Yandex.
In Vireo
Excluded, reported under Search engines

What it does

Yandex's main indexing robot.

Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) The main indexing robot.

Yandex

User agent

Requests from YandexBot carry a user agent like this:

Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots)

Device Detector identifies it with this pattern (a case-insensitive regular expression):

(Yandex(?:(?:\.Gazeta |Accessibility|Com|Mobile|MobileScreenShot|RenderResources|Screenshot|Sprav)?Bot|(?:AdNet|Antivirus|Blogs|Calendar|Catalog|Dialogs|Direct(?:Dyn)?|Favicons|ForDomain|ImageResizer|Images|Market|Media(?:naBot)?|Metrika|News(?:links)?|OntoDB(?:API)?|Pagechecker|Partner|RCA|SearchShop|(?:News|Site)links|Tracker|Turbo|Userproxy|Verticals|Vertis|Video(?:Parser)?|Webmaster))|YaDirectFetcher)

Anyone can put any name in a user agent. Where the operator publishes IP ranges, checking the request's address against them is the way to confirm a visit is genuine.

How to block or allow YandexBot

Add a group for its token to the robots.txt file at the root of your site. To keep it out of the whole site:

User-agent: YandexBot
Disallow: /

To let it crawl everything, use an empty Disallow: or Allow: / instead. WordPress serves a virtual robots.txt that many SEO plugins let you edit. A real robots.txt file uploaded to the site root takes its place.

User-agent: YandexBot # will be used only by the main indexing bot

Yandex

Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) The main indexing robot. Yes

Yandex

robots.txt is a request, so it only works on crawlers that honour it. A bot that ignores it has to be blocked at the server or firewall, by user agent or IP address.

Worth knowing

The quoted 'Yes' is the 'Takes into account the General rules specified in robots.txt' column. Yandex doesn't publish its crawler IPs, so verify with reverse DNS (yandex.ru, yandex.net or yandex.com). Some other Yandex robots ignore rules aimed at User-agent: *.

To avoid unintentional blocking by site owners, they may ignore the file's restrictive directives robots.txt designed for arbitrary robots ( User-agent: *).

Yandex

How Vireo counts YandexBot

Vireo keeps requests from YandexBot out of your pageviews and visitors and counts them on its bot report under Search engines, so you can see how much it left out rather than wonder why your server logs show more traffic.

Checked by running Vireo Analytics' own bot filter over YandexBot's user agents.

Sources

More search engines

All search engines

See the bots on your own site

Vireo Analytics is a free WordPress plugin. It keeps over 800 known crawlers out of your numbers and shows you what it filtered, by category.

Get Vireo free on WordPress.org