ccBot crawler

ccBot crawler is a crawler run by Common Crawl.

Operator
Common Crawl
Type
Crawler
Official page
commoncrawl.org
robots.txt token
CCBot
Follows robots.txt
Yes, according to Common Crawl.
Runs JavaScript
No, according to Common Crawl.
Published IP ranges
index.commoncrawl.org/ccbot.json
In Vireo
Excluded, reported under AI crawlers

What it does

Builds Common Crawl's open repository of web crawl data, which anyone can download and analyze.

producing and maintaining an open repository of web crawl data that is universally accessible and analyzable by anyone.

Common Crawl

Crawl frequency, per ai.robots.txt: Monthly at present.

User agent

Requests from ccBot crawler carry a user agent like this:

CCBot/2.0 (http://commoncrawl.org/faq/)

Device Detector identifies it with this pattern (a case-insensitive regular expression):

CCBot

Anyone can put any name in a user agent. Where the operator publishes IP ranges, checking the request's address against them is the way to confirm a visit is genuine.

How to block or allow ccBot crawler

Add a group for its token to the robots.txt file at the root of your site. To keep it out of the whole site:

User-agent: CCBot
Disallow: /

To let it crawl everything, use an empty Disallow: or Allow: / instead. WordPress serves a virtual robots.txt that many SEO plugins let you edit. A real robots.txt file uploaded to the site root takes its place.

To prevent Common Crawl from crawling your website, include the following in your robots.txt: User-agent: CCBot Disallow: /

Common Crawl

CCBot is an automated crawler, checking first the robots.txt, and if crawling a page is allowed, fetches pages using HTTP GET requests.

Common Crawl

robots.txt is a request, so it only works on crawlers that honour it. A bot that ignores it has to be blocked at the server or firewall, by user agent or IP address.

Worth knowing

Common Crawl says other crawlers fake its user agent. Check reverse DNS ends in crawl.commoncrawl.org.

Please note that we are aware of crawlers falsely identifying themselves as CCBot.

Common Crawl

Does ccBot crawler run JavaScript?

No. Common Crawl says it doesn't.

Currently, JavaScript is not executed and Cookies are not used.

Common Crawl

How Vireo counts ccBot crawler

Vireo keeps requests from ccBot crawler out of your pageviews and visitors and counts them on its bot report under AI crawlers, so you can see how much it left out rather than wonder why your server logs show more traffic.

Because it's an AI crawler, its requests also count toward the pages AI crawlers took, which Vireo shows next to the readers AI assistants sent back. See how the AI assistants card works.

Checked by running Vireo Analytics' own bot filter over ccBot crawler's user agents.

Sources

More ai crawlers

All ai crawlers

See the bots on your own site

Vireo Analytics is a free WordPress plugin. It keeps over 800 known crawlers out of your numbers and shows you what it filtered, by category.

Get Vireo free on WordPress.org