Common Crawl

AI training

CCBot

Builds Common Crawl's open archive of the web, which many AI labs use as training data.

Follows robots.txt
User agent token
CCBot
Operator
Common Crawl
Verification
Published IP list , reverse DNS crawl.commoncrawl.org

What an AI training crawler does

A training crawler gathers public pages that may end up in the data used to train AI models. It sends no visitors back, so blocking it is the usual choice for sites that don't want their content used that way.

Common Crawl warns that other bots pretend to be CCBot, so check the IP.

Block or allow CCBot in robots.txt

Add one of these to the robots.txt file at the root of your domain. Crawlers read it before they fetch anything else.

Block CCBot

User-agent: CCBot
Disallow: /

Allow CCBot

User-agent: CCBot
Allow: /
Browse all 66 crawlersAI and search bots from 16 companies, with the rules to block or allow each one.

Questions, answered.

CCBot is an AI training crawler run by Common Crawl. Builds Common Crawl's open archive of the web, which many AI labs use as training data.

Yes. Common Crawl says CCBot follows robots.txt. Common Crawl warns that other bots pretend to be CCBot, so check the IP.

Add "User-agent: CCBot" and "Disallow: /" to your robots.txt.

Check the request's IP address against the list Common Crawl publishes. A user agent alone proves nothing, since any script can copy it.

No. CCBot collects pages for model training and sends no visitors back.

Block it if you don't want your pages used to train models. It sends no visitors, so blocking it costs you no traffic.

In its own documentation at https://commoncrawl.org/ccbot. Every fact on this page comes from there.

Still have questions? We're happy to help.

See which bots really read your site.

NoirTrack counts CCBot and every other bot apart from real visitors, and its firewall can block them before they reach your pages.

Start free trial

Free 14-day trial, no credit card. The server SDK also sees bots that never run JavaScript.