# Bot filtering

> Bot filtering is separating automated requests from human ones in analytics. The important design decision is where it happens: filtering at ingest discards the request permanently, while filtering at read time keeps it and decides what to show — which is the difference between a number you can revise and one you cannot.

**Category:** Analytics
**Also known as:** bot traffic filtering, spider filtering
**Updated:** 2026-09-09
**Source:** https://truestat.io/glossary/bot-filtering

---

## Ingest versus read time

Filtering at ingest means the request is classified as a bot and thrown away. It is cheaper, and it is permanent: when the classifier turns out to have been wrong, or when a crawler that was unknown last month gets a name this month, the history cannot be recovered.

Filtering at read time means everything is stored with a label and the label decides what a given view counts. It costs storage and a little query complexity, and in exchange every past decision stays revisable. A crawler recognised today can be counted correctly in last year's data.

For a product whose whole subject is AI crawler traffic, discarding bot requests at ingest would throw away the feature.

## The categories worth keeping apart

"Bot" is too coarse to be useful. A training crawler, an assistant fetching an answer, a search indexer, an uptime monitor and a link-preview fetcher are five different things, and folding them together produces a number that means nothing.

The split that matters most in practice is preview bots. `facebookexternalhit`, `Twitterbot` and `Slackbot` fetch a page to draw a link card, and every share produces one — so a popular post can generate more preview fetches than a site has real AI visits. Counting them as AI inflates the figure visibly.

## What a good bot report looks like

Four properties, and most analytics products have none of them because bot traffic is treated as a nuisance to be removed rather than a subject.

**Separate axes.** Human traffic and bot traffic on the same chart makes a training sweep look like a launch. They answer different questions and belong in different series.

**Named operators, not a lump.** "Bots: 4,312" tells a reader nothing they can act on. "OpenAI 900, Meta 2,100, unidentified 1,312" tells them where the bandwidth went and which of it was worth having.

**Confidence shown, not implied.** A verified count and an unverified count are different claims, and merging them means the number cannot survive the question of how you know.

**Reversible classification.** The list of known crawlers grows monthly. A product that labels at read time can apply today's knowledge to last year's traffic; one that filtered at ingest cannot, and its history is permanently as good as its classifier was on the day.

## What goes wrong

- Filtering at ingest. It is the one decision here that cannot be undone.
- Counting link previews as AI traffic.
- Missing your own infrastructure. A favicon fetcher, an uptime monitor or a preview service you run yourself will show up as crawler traffic on your customers' sites.
- Assuming an unrecognised bot is a human. It is the safer default, because it understates the AI figure rather than inflating it, and for a number sold on honesty that is the correct direction to be wrong — but it is still a known error, and the size of it is unknowable by definition.

## Worth doing

- Store the classification alongside the event rather than instead of it.
- Keep a separate axis for bot traffic in every chart. A crawl is not a visit and should not share a line with one.

## Questions

### Does Google Analytics filter bots?

It filters known bots from a list it maintains, and it never sees the crawlers discussed here at all — they do not run JavaScript, so no event is ever sent. The filtering question only arises for bots that execute scripts.

### Should bot traffic count toward my plan usage?

That is a policy decision rather than a technical one, and it deserves to be stated plainly wherever the number is shown. Charging for traffic a customer did not ask for and cannot see is a way to lose them.

### Can I tell a headless browser from a real one?

Sometimes, and never reliably. `HeadlessChrome` in the user agent is an admission rather than a detection, and anything determined to hide will remove it. Behavioural signals exist but they misclassify real people — screen readers, old devices, privacy browsers — which is a worse error than missing a bot.

## Related

- https://truestat.io/glossary/server-side-tracking
- https://truestat.io/glossary/user-agent
- https://truestat.io/glossary/spoofed-crawler

---

From the TrueStat glossary — https://truestat.io/glossary. Privacy-first web analytics that also shows you which AI assistants are reading your site.
