CCBot

Honors robots.txt

CCBot crawls the web to build the open, freely-downloadable Common Crawl dataset, which many AI labs use as raw pretraining data, either directly or via datasets derived from it.

User-agent
CCBot
Operator
Common Crawl (nonprofit)

Allow in robots.txt

User-agent: CCBot
Allow: /

Block in robots.txt

User-agent: CCBot
Disallow: /

CCBot isn't tied to one named AI vendor — blocking it reduces your presence in a dataset used by many different labs and open-source model builders, not just one company.

Frequently asked questions

No — Common Crawl is an independent nonprofit dataset used as training input by numerous AI labs, so blocking or allowing CCBot is a broader decision than opting in or out of any single vendor.

Verify your robots.txt against every AI crawler

Scoutern's free AI bot checker tests exactly which bots can reach your site.