Common Crawl
CCBot
Builds the open web dataset most model builders start from.
When you see it: Common Crawl trains nothing itself, but its archive is a training input for most of the models on this page — so it belongs here.
How we check it is really Common Crawl
- 01
The user agent has to contain CCBot.
- 02
The address is compared against Common Crawl's published ranges, which we refetch every six hours.
Common Crawl publishes no reverse DNS record, so an address outside the list is recorded as unverified rather than as a forgery — the list may simply be behind.
What each result means
- verified
- The address matched Common Crawl's own published records. This really was them.
- unverified
- We could not confirm it, and that is all it means. The operator may publish nothing to check against, a published list may be behind the addresses actually in use, or DNS may simply have been slow. It is not an accusation.
- spoofed
- The address belongs to somebody else — its own reverse DNS record resolves to a different owner. That is positive evidence of a forgery rather than a failure to find evidence, which is why it is reported as its own result and never folded into the row above.
The distinction costs nothing to make and is the reason these numbers can be shown to someone who will ask where they came from. A report that cannot separate “we do not know” from “this was faked” invites exactly one question it cannot answer.
The lists this is checked against
Published by Common Crawl and refetched every six hours. Open them; they are the same files we read.
- Operator
- Common Crawl
- Token
- CCBot
- Counted as
- AI training
- Verification
- Published IP ranges
- Runs JavaScript
- No
- Operator docs
- Published
User agent
CCBotMatched as a token, not as a full browser string. Common Crawl adds and changes the version and product text around it without notice, so a filter pinned to the whole header stops working the first time they bump a version while one matching CCBot keeps working. The same string is what Common Crawl names in robots.txt.
One real request
CCBot/2.0 (https://commoncrawl.org/faq/)Your analytics never mentioned CCBot because it cannot see it.
This crawler takes your HTML and leaves without running a line of JavaScript, so a browser tag records nothing. TrueStat reads it server-side, checks the address, and shows you which pages it fetched — including the ones that returned a 404.