Diffbot
Diffbot is an AI data scraper run by Diffbot.
- Operator
- Diffbot
- Type
- AI Data Scraper
- Official page
- diffbot.com
- robots.txt token
Diffbot- Follows robots.txt
- Partly, according to Diffbot.
- In Vireo
- Excluded, reported under AI crawlers
What it does
Crawls the web for Diffbot's Knowledge Graph and web search. Diffbot says it isn't used for AI training.
General, proactive web crawling for building a general search engine. This allows websites to be discovered and cited from the Diffbot Knowledge Graph and web search services in response to keyword queries. It is not used for AI training.
Diffbot
User agent
Requests from Diffbot carry user agents like these:
Snap/1.0 (+https://www.diffbot.com/dev/docs/crawl/)Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.9.1.2) Gecko/20090729 Firefox/3.5.2 (.NET CLR 3.5.30729; Diffbot/0.1; +http://www.diffbot.com)Device Detector identifies it with this pattern (a case-insensitive regular expression):
.+diffbotAnyone can put any name in a user agent. Where the operator publishes IP ranges, checking the request's address against them is the way to confirm a visit is genuine.
How to block or allow Diffbot
Add a group for its token to the robots.txt file at the root of your site. To keep it out of the whole site:
User-agent: Diffbot
Disallow: /To let it crawl everything, use an empty Disallow: or Allow: / instead. WordPress serves a virtual robots.txt that many SEO plugins let you edit. A real robots.txt file uploaded to the site root takes its place.
To whitelist Diffbot for a site, specify the appropriate Diffbot user-agents in the site’s robots.txt.
Diffbot
By default Diffbot's web crawls adhere to a site’s robots.txt instructions, including the disallow and crawl-delay directives. In specific cases — typically because of a partnership or agreement you have with the site to be crawled — the robots.txt instruction can be ignored/overridden.
Diffbot
robots.txt is a request, so it only works on crawlers that honour it. A bot that ignores it has to be blocked at the server or firewall, by user agent or IP address.
Worth knowing
Diffbot customers running Crawlbot or Extract set their own crawl settings and can set a custom user agent. A separate Diffbot-User agent acts on behalf of a human user.
When users use Extract or Crawlbot APIs, they are defining their own crawl parameters using hosted software.
Diffbot
How Vireo counts Diffbot
Vireo keeps requests from Diffbot out of your pageviews and visitors and counts them on its bot report under AI crawlers, so you can see how much it left out rather than wonder why your server logs show more traffic.
Because it's an AI crawler, its requests also count toward the pages AI crawlers took, which Vireo shows next to the readers AI assistants sent back. See how the AI assistants card works.
Checked by running Vireo Analytics' own bot filter over Diffbot's user agents.
Sources
- Diffbot: crawler documentation (checked October 2026)
- Diffbot information page (docs.diffbot.com)
- Matomo Device Detector, bots.yml (LGPL-3.0)
- ai.robots.txt (MIT)
Other bots from Diffbot
More ai crawlers
- Ai2Bot (Allen Institute for AI (Ai2))
- Ai2Bot-DeepResearchEval (The Allen Institute for Artificial Intelligence)
- Ai2Bot-Dolma (The Allen Institute for Artificial Intelligence)
- Amazonbot (Amazon)
- Amazonbot-Video (Amazon.com, Inc.)
- AmazonBuyForMe (Amazon.com, Inc.)
- Amzn-SearchBot (Amazon.com, Inc.)
- Amzn-User (Amazon.com, Inc.)
- Andibot (Andi)
- Anthropic AI (Anthropic, PBC)
See the bots on your own site
Vireo Analytics is a free WordPress plugin. It keeps over 800 known crawlers out of your numbers and shows you what it filtered, by category.
Get Vireo free on WordPress.org