# Doubaobot

> Doubaobot is ByteDance's crawler for traffic that is neither answering nor training. ByteDance's crawler for the Doubao assistant. ByteDance fetched your page for Doubao, its consumer assistant. Written both `DoubaoBot` and `Doubaobot` in the wild; both are matched.

**Category:** AI crawlers
**Also known as:** ByteDance Doubaobot
**Updated:** 2026-09-09
**Source:** https://truestat.io/glossary/bytedance-doubaobot

---

## What Doubaobot is

ByteDance publishes no crawler documentation and no address ranges, and its main crawler has a long-standing reputation for ignoring robots.txt. What is known about these tokens comes from site owners' logs rather than from the company.

On this site Doubaobot is counted as other AI traffic, which is the row it appears in on your dashboard. The category matters because it decides which number moves when this crawler visits, and the four are reported separately for exactly that reason.

## What its traffic looks like

Irregular and low volume. These fetches happen for a specific reason — an ad check, a preview, an infrastructure request — rather than on a schedule, so there is no pattern worth reading into unless the count is unexpectedly high.

## How it is verified

ByteDance publishes neither a machine-readable list of the addresses this crawler uses nor a reverse DNS record pointing back at itself. The user-agent header is therefore the only evidence, and anyone can send it — so a visit is recorded as an unverified claim rather than as a confirmed fetch.

That is not a gap in the measurement so much as an honest ceiling. A report that showed these as verified would be inventing a confidence nobody can support, and the number it produced would be the first thing to fall apart when somebody asked where it came from.

## If you block it

Very little. This fetch is neither an answer nor training, so blocking it costs you nothing except the record of ByteDance having visited.

There is not much to decide here. This fetch is neither answering a question nor collecting training material, so it neither brings readers nor takes anything you would miss.

It is worth watching the volume rather than the intent. If a crawler in this category becomes a large share of your requests, that is a bandwidth question, and it is answered with a rate limit rather than a robots rule.

## How it differs from ByteDance's other crawlers

ByteDance sends 3 crawlers and gives each its own token, so they can be allowed and refused separately. The tokens differ by a few characters and what the fetch means does not:

TikTokSpider — ai training. ByteDance's crawler. No operator documentation exists for it.
Bytespider — ai training. ByteDance's training crawler.

Confusing two of these is the most common mistake with this operator, and it is expensive in one direction: a rule meant for the training crawler that lands on the search crawler removes you from results without stopping any training.

## Questions

### Is Doubaobot in my logs really ByteDance?

There is no way to prove it. ByteDance publishes neither an address list nor a reverse DNS record for this crawler, so the user agent is the only evidence and anyone can send it. We record the visit as an unverified claim rather than presenting it as confirmed.

### How do I tell Doubaobot apart from ByteDance's other crawlers?

By the token, and only by the token. ByteDance also sends TikTokSpider, Bytespider, and the strings differ by a few characters while the jobs differ completely — one may be answering somebody's question while another collects training material. Match on the exact token rather than on the operator's name appearing anywhere in the header, or the four end up counted as one.

### Does Doubaobot run JavaScript?

No. It requests the HTML and leaves. That is why a browser-based analytics tag never records these visits: the tag is JavaScript, and nothing runs it. Seeing this crawler at all requires reading it server-side.

### Why is Doubaobot counted as Other AI traffic and not something else?

Because that is the job this particular crawler does. ByteDance runs more than one — TikTokSpider, Bytespider — and they are counted separately because the decisions about them are separate. Filing them together under one operator name would let ordinary indexing look like AI interest, which is the single most misleading thing an AI traffic report can do.

### Why does Doubaobot not show up in Google Analytics?

Google Analytics runs in the browser. This crawler never opens a browser — it requests the HTML, reads it, and leaves, so the tracking script is never executed and no event is ever sent. Every browser-based analytics tool has the same blind spot, which is why crawler traffic has to be read from the server side to be seen at all.

### Why can a Doubaobot visit not be verified?

ByteDance publishes neither a machine-readable list of the addresses its crawler uses nor a reverse DNS record pointing back at itself. With neither, the user-agent header is the only evidence, and anyone can send it. Rather than showing a confidence we cannot support, the visit is recorded as an unverified claim — which is also the honest answer to give anyone asking how much of your AI traffic is real.

## Related

- https://truestat.io/glossary/bytedance-tiktokspider
- https://truestat.io/glossary/bytedance-bytespider
- https://truestat.io/glossary/fcrdns
- https://truestat.io/glossary/spoofed-crawler

---

From the TrueStat glossary — https://truestat.io/glossary. Privacy-first web analytics that also shows you which AI assistants are reading your site.
