Baidu

Baiduspider

IndexingIP ranges + reverse DNS

Baidu's search crawler.

When you see it: Ordinary Baidu indexing.

How we check it is really Baidu

  1. 01

    The user agent has to contain Baiduspider.

  2. 02

    The address is compared against Baidu's published ranges, which we refetch every six hours.

  3. 03

    If the ranges do not settle it, the address's reverse DNS record is resolved and then resolved forward again, so a faked record cannot pass.

A visit that matches neither is recorded as unverified, not as a forgery. A range list can be incomplete, and an accusation needs evidence.

What each result means

verified
The address matched Baidu's own published records. This really was them.
unverified
We could not confirm it, and that is all it means. The operator may publish nothing to check against, a published list may be behind the addresses actually in use, or DNS may simply have been slow. It is not an accusation.
spoofed
The address belongs to somebody else — its own reverse DNS record resolves to a different owner. That is positive evidence of a forgery rather than a failure to find evidence, which is why it is reported as its own result and never folded into the row above.

The distinction costs nothing to make and is the reason these numbers can be shown to someone who will ask where they came from. A report that cannot separate “we do not know” from “this was faked” invites exactly one question it cannot answer.

No list to check against

Baidu publishes no machine-readable address file for this crawler, so verification runs through reverse DNS instead — its records resolve under .crawl.baidu.com, .baidu.com.

Operator
Baidu
Token
Baiduspider
Counted as
Search indexing
Verification
IP ranges + reverse DNS
Runs JavaScript
No
Operator docs
Published

User agent

Baiduspider

Matched as a token, not as a full browser string. Baidu adds and changes the version and product text around it without notice, so a filter pinned to the whole header stops working the first time they bump a version while one matching Baiduspider keeps working. The same string is what Baidu names in robots.txt.

Baidu's other crawlers

Baidu sends 2 crawlers, each with its own token and its own job. They are counted separately for that reason — the tokens differ by a few characters and what the fetch means does not, so folding them into one operator row would let indexing look like AI interest.

TokenCounted asVerificationWhat it fetches for
Baiduspiderthis pageIndexingIP ranges + reverse DNSBaidu's search crawler.
ERNIEBotTrainingIP ranges + reverse DNSBaidu's training crawler for ERNIE.

Your analytics never mentioned Baiduspider because it cannot see it.

This crawler takes your HTML and leaves without running a line of JavaScript, so a browser tag records nothing. TrueStat reads it server-side, checks the address, and shows you which pages it fetched — including the ones that returned a 404.