Scrapy

Scrapy is a crawler run by Zyte.

Operator
Zyte
Type
Crawler
Official page
scrapy.org
robots.txt token
Scrapy
In Vireo
Excluded, reported under Scripts and clients

What it does

"AI and machine learning applications often need large amounts of quality data, and web data extraction is a fast, efficient way to build structured data sets."

ai.robots.txt describes its function as: Scrapes data for a variety of uses including training AI.

Source: ai.robots.txt.

User agent

Requests from Scrapy carry a user agent like this:

Scrapy/1.0.3.post6+g2d688cd (+http://scrapy.org)

Device Detector identifies it with this pattern (a case-insensitive regular expression):

Scrapy

Anyone can put any name in a user agent. Where the operator publishes IP ranges, checking the request's address against them is the way to confirm a visit is genuine.

How to block or allow Scrapy

Add a group for its token to the robots.txt file at the root of your site. To keep it out of the whole site:

User-agent: Scrapy
Disallow: /

To let it crawl everything, use an empty Disallow: or Allow: / instead. WordPress serves a virtual robots.txt that many SEO plugins let you edit. A real robots.txt file uploaded to the site root takes its place.

robots.txt is a request, so it only works on crawlers that honour it. A bot that ignores it has to be blocked at the server or firewall, by user agent or IP address.

How Vireo counts Scrapy

Vireo keeps requests from Scrapy out of your pageviews and visitors and counts them on its bot report under Scripts and clients, so you can see how much it left out rather than wonder why your server logs show more traffic.

Checked by running Vireo Analytics' own bot filter over Scrapy's user agents.

Sources

See the bots on your own site

Vireo Analytics is a free WordPress plugin. It keeps over 800 known crawlers out of your numbers and shows you what it filtered, by category.

Get Vireo free on WordPress.org