Should you block CCBot? Common Crawl's dataset crawler

CCBot is operated by Common Crawl. It builds a public dataset that other systems, including AI trainers, consume downstream.

Who should block CCBot, and who should not?

Block it if you do not want future pages in open web datasets. Common Crawl publishes its archive for anyone to use, including AI developers who never crawl you themselves, so this one rule reaches further than a block on any single AI company's crawler. The block covers future crawls; for pages already collected, Common Crawl's route is a legal opt-out request. The cost to readers is nil; the loss is research and archive presence. Verify requests before judging its behaviour from logs, because Common Crawl says impostors use its name.

The verified facts

User-agent tokenCCBot/2.0
Full user agent in logsCCBot/2.0 (https://commoncrawl.org/faq/)
OperatorCommon Crawl
PurposeArchive and dataset crawlers
Honors robots.txtYes, per the operator's documentation.
Published IP rangesindex.commoncrawl.org/ccbot.json
Reverse DNS ends incrawl.commoncrawl.org

What does Common Crawl’s documentation add?

  • Common Crawl says it knows of crawlers falsely identifying themselves as CCBot, and recommends verifying requests. Reverse DNS works over IPv4; it is not yet supported over IPv6, where the IP file is the check.
  • It obeys Crawl-delay, slows down on HTTP 429 or 5xx responses, and by default waits a few seconds between requests to the same site.
  • Common Crawl says its dataset is a random sample of the web, not a full archive of any site. It does not execute JavaScript or use cookies.
  • A robots.txt block stops future crawls. Common Crawl does not document removing pages from past crawls; for material already collected, its route is a legal opt-out request to info@commoncrawl.org, which it lists in its public Opt-Out Ledger to alert downstream users.

How do you block CCBot?

Add this to your robots.txt:

User-agent: CCBot
Disallow: /

Once you give a crawler its own group, it stops reading your User-agent: * rules, so run the whole file through the AI crawler robots.txt tester before you deploy it.

How do you tell a real CCBot request from a fake one?

A user-agent match proves nothing, since any client can send CCBot/2.0. Check the IP against the published range file above, re-fetched on a schedule. Or reverse-resolve the IP, check the name ends in crawl.commoncrawl.org, and forward-resolve that name back to the same IP. A request that fails is not Common Crawl. The log verification guide has a script for the IP check and a worked reverse-DNS example, or paste the IP into the AI crawler IP verifier, which runs these checks against Common Crawl’s current data.

What does blocking CCBot cost you?

Content excluded from Common Crawl archives and public web datasets.