ccBot crawler
ccBot crawler is a crawler run by Common Crawl.
- Operator
- Common Crawl
- Type
- Crawler
- Official page
- commoncrawl.org
- robots.txt token
CCBot- Follows robots.txt
- Yes, according to Common Crawl.
- Runs JavaScript
- No, according to Common Crawl.
- Published IP ranges
- index.commoncrawl.org/ccbot.json
- In Vireo
- Excluded, reported under AI crawlers
What it does
Builds Common Crawl's open repository of web crawl data, which anyone can download and analyze.
producing and maintaining an open repository of web crawl data that is universally accessible and analyzable by anyone.
Common Crawl
Crawl frequency, per ai.robots.txt: Monthly at present.
User agent
Requests from ccBot crawler carry a user agent like this:
CCBot/2.0 (http://commoncrawl.org/faq/)Device Detector identifies it with this pattern (a case-insensitive regular expression):
CCBotAnyone can put any name in a user agent. Where the operator publishes IP ranges, checking the request's address against them is the way to confirm a visit is genuine.
How to block or allow ccBot crawler
Add a group for its token to the robots.txt file at the root of your site. To keep it out of the whole site:
User-agent: CCBot
Disallow: /To let it crawl everything, use an empty Disallow: or Allow: / instead. WordPress serves a virtual robots.txt that many SEO plugins let you edit. A real robots.txt file uploaded to the site root takes its place.
To prevent Common Crawl from crawling your website, include the following in your robots.txt: User-agent: CCBot Disallow: /
Common Crawl
CCBot is an automated crawler, checking first the robots.txt, and if crawling a page is allowed, fetches pages using HTTP GET requests.
Common Crawl
robots.txt is a request, so it only works on crawlers that honour it. A bot that ignores it has to be blocked at the server or firewall, by user agent or IP address.
Worth knowing
Common Crawl says other crawlers fake its user agent. Check reverse DNS ends in crawl.commoncrawl.org.
Please note that we are aware of crawlers falsely identifying themselves as CCBot.
Common Crawl
Does ccBot crawler run JavaScript?
No. Common Crawl says it doesn't.
Currently, JavaScript is not executed and Cookies are not used.
Common Crawl
How Vireo counts ccBot crawler
Vireo keeps requests from ccBot crawler out of your pageviews and visitors and counts them on its bot report under AI crawlers, so you can see how much it left out rather than wonder why your server logs show more traffic.
Because it's an AI crawler, its requests also count toward the pages AI crawlers took, which Vireo shows next to the readers AI assistants sent back. See how the AI assistants card works.
Checked by running Vireo Analytics' own bot filter over ccBot crawler's user agents.
Sources
- Common Crawl: crawler documentation (checked October 2026)
- ccBot crawler information page (commoncrawl.org)
- Matomo Device Detector, bots.yml (LGPL-3.0)
- ai.robots.txt (MIT)
More ai crawlers
- Ai2Bot (Allen Institute for AI (Ai2))
- Ai2Bot-DeepResearchEval (The Allen Institute for Artificial Intelligence)
- Ai2Bot-Dolma (The Allen Institute for Artificial Intelligence)
- Amazonbot (Amazon)
- Amazonbot-Video (Amazon.com, Inc.)
- AmazonBuyForMe (Amazon.com, Inc.)
- Amzn-SearchBot (Amazon.com, Inc.)
- Amzn-User (Amazon.com, Inc.)
- Andibot (Andi)
- Anthropic AI (Anthropic, PBC)
See the bots on your own site
Vireo Analytics is a free WordPress plugin. It keeps over 800 known crawlers out of your numbers and shows you what it filtered, by category.
Get Vireo free on WordPress.org