Common Crawl

CCBot

TrainingPublished IP ranges

Builds the open web dataset most model builders start from.

When you see it: Common Crawl trains nothing itself, but its archive is a training input for most of the models on this page — so it belongs here.

How we check it is really Common Crawl

  1. 01

    The user agent has to contain CCBot.

  2. 02

    The address is compared against Common Crawl's published ranges, which we refetch every six hours.

Common Crawl publishes no reverse DNS record, so an address outside the list is recorded as unverified rather than as a forgery — the list may simply be behind.

What each result means

verified
The address matched Common Crawl's own published records. This really was them.
unverified
We could not confirm it, and that is all it means. The operator may publish nothing to check against, a published list may be behind the addresses actually in use, or DNS may simply have been slow. It is not an accusation.
spoofed
The address belongs to somebody else — its own reverse DNS record resolves to a different owner. That is positive evidence of a forgery rather than a failure to find evidence, which is why it is reported as its own result and never folded into the row above.

The distinction costs nothing to make and is the reason these numbers can be shown to someone who will ask where they came from. A report that cannot separate “we do not know” from “this was faked” invites exactly one question it cannot answer.

The lists this is checked against

Published by Common Crawl and refetched every six hours. Open them; they are the same files we read.

Operator
Common Crawl
Token
CCBot
Counted as
AI training
Verification
Published IP ranges
Runs JavaScript
No
Operator docs
Published

User agent

CCBot

Matched as a token, not as a full browser string. Common Crawl adds and changes the version and product text around it without notice, so a filter pinned to the whole header stops working the first time they bump a version while one matching CCBot keeps working. The same string is what Common Crawl names in robots.txt.

One real request

CCBot/2.0 (https://commoncrawl.org/faq/)

Your analytics never mentioned CCBot because it cannot see it.

This crawler takes your HTML and leaves without running a line of JavaScript, so a browser tag records nothing. TrueStat reads it server-side, checks the address, and shows you which pages it fetched — including the ones that returned a 404.