ByteDance

Bytespider

TrainingUser agent only

ByteDance's training crawler.

When you see it: ByteDance is gathering training data.

How we check it is really ByteDance

  1. 01

    The user agent has to contain Bytespider.

ByteDance publishes neither IP ranges nor a reverse DNS record, so this is the honest ceiling: the visit is recorded as a claim. We would rather show you an unverified row than a tick we cannot back.

What each result means

verified
The address matched ByteDance's own published records. This really was them.
unverified
We could not confirm it, and that is all it means. The operator may publish nothing to check against, a published list may be behind the addresses actually in use, or DNS may simply have been slow. It is not an accusation.
spoofed
The address belongs to somebody else — its own reverse DNS record resolves to a different owner. That is positive evidence of a forgery rather than a failure to find evidence, which is why it is reported as its own result and never folded into the row above.

The distinction costs nothing to make and is the reason these numbers can be shown to someone who will ask where they came from. A report that cannot separate “we do not know” from “this was faked” invites exactly one question it cannot answer.

No list to check against

ByteDance publishes neither an address file nor a reverse DNS record for this crawler. There is nothing to check a request against, which is why every visit from it is recorded as an unverified claim.

Operator
ByteDance
Token
Bytespider
Counted as
AI training
Verification
User agent only
Runs JavaScript
No
Operator docs
None published

ByteDance publishes no crawler documentation. This token is recognised from observed traffic and community lists, so treat the classification as our reading rather than the operator's statement.

User agent

Bytespider

Matched as a token, not as a full browser string. ByteDance adds and changes the version and product text around it without notice, so a filter pinned to the whole header stops working the first time they bump a version while one matching Bytespider keeps working. The same string is what ByteDance names in robots.txt.

Site owners consistently report this crawler fetching pages that robots.txt disallows, including paths no link points at. Worth knowing when you read its numbers: a robots rule is not evidence these requests stopped.

One real request

Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)

ByteDance's other crawlers

ByteDance sends 3 crawlers, each with its own token and its own job. They are counted separately for that reason — the tokens differ by a few characters and what the fetch means does not, so folding them into one operator row would let indexing look like AI interest.

TokenCounted asVerificationWhat it fetches for
Bytespiderthis pageTrainingUser agent onlyByteDance's training crawler.
TikTokSpiderTrainingUser agent onlyByteDance's crawler. No operator documentation exists for it.
DoubaobotOther AI trafficUser agent onlyByteDance's crawler for the Doubao assistant.

Your analytics never mentioned Bytespider because it cannot see it.

This crawler takes your HTML and leaves without running a line of JavaScript, so a browser tag records nothing. TrueStat reads it server-side, checks the address, and shows you which pages it fetched — including the ones that returned a 404.