CCBot: user agent, IP verification and robots.txt
CCBot is operated by Common Crawl Foundation. It is classified here as Training (Open crawl archive). Verification method: Reverse DNS only.
Exact user-agent string
CCBot/2.0 (https://commoncrawl.org/faq/)
robots.txt token: CCBot. The user-agent string is self-asserted and can be
forged by anyone. It is a claim, not proof.
How to verify CCBot is genuine
Verify by reverse DNS. Resolve the source IP to a hostname, confirm it belongs to Common Crawl Foundation, then resolve that hostname forward and confirm it returns the original IP. A reverse lookup alone is not sufficient.
How to block or allow CCBot
User-agent: CCBot
Disallow: / Replace Disallow with Allow to permit it explicitly.
What blocking CCBot costs you
CCBot feeds an open, freely redistributable archive. Blocking it removes you from a dataset that many downstream model builders use, which is a broader training opt-out than blocking any single operator, but it also removes you from academic and non-profit research corpora that have nothing to do with commercial AI.
What is specific to CCBot
Common Crawl is the only operator in this reference that publicly acknowledges being impersonated. Its own CCBot page states that it is aware of crawlers falsely identifying themselves as CCBot and recommends verifying user-agent strings to ensure authenticity. It moved CCBot onto dedicated IP ranges with reverse DNS specifically so this could be checked: a genuine request resolves under crawl.commoncrawl.org, and the forward lookup returns the same address. The exception is IPv6, where Common Crawl states reverse DNS is not yet supported, so an IPv6 CCBot request has no verification path at all. Blocking is also unusually consequential in one direction: content already collected sits in published archives, so a robots.txt change stops future collection but does not retract past crawls.
Verify a request claiming to be CCBot
Check a CCBot request against Common Crawl Foundation's published ranges
Source
Primary documentation: https://commoncrawl.org/ccbot. Last verified 2026-08-02.