Crawler directory

Every crawler we can name, and how we know it is really them.

A user agent is a claim, not proof — anyone can send curl -A "GPTBot". Each of these is checked against the operator's own published address ranges or its reverse DNS record, and where neither exists the crawler's page says so instead of implying a check we cannot perform.

Crawlers
59
Operators
23
Verifiable
35
Buckets
4

Plus 1,500 generic patterns for scrapers, monitors and headless browsers — counted in your dashboard, not named here.

23 operators, 13 of them publishing something we can check.

An operator earns a name here by being classified in the detector, not by being famous. The ones that publish nothing still get a row — they simply cannot get a tick.

  • Google

    7 crawlers · verifiable

  • OpenAI

    5 crawlers · verifiable

  • xAI

    5 crawlers

  • Anthropic

    4 crawlers · verifiable

  • Meta

    4 crawlers

  • Alibaba

    3 crawlers

  • Amazon

    3 crawlers

  • ByteDance

    3 crawlers

  • Microsoft Bing

    3 crawlers · verifiable

  • Mistral

    3 crawlers · verifiable

  • Moonshot AI

    3 crawlers · verifiable

  • Baidu

    2 crawlers · verifiable

  • DuckDuckGo

    2 crawlers · verifiable

  • Huawei

    2 crawlers · verifiable

  • Perplexity

    2 crawlers · verifiable

  • Allen AI

    1 crawler

  • Apple

    1 crawler · verifiable

  • Cohere

    1 crawler

  • Common Crawl

    1 crawler · verifiable

  • DeepSeek

    1 crawler

  • Yandex

    1 crawler · verifiable

  • You.com

    1 crawler

  • Zhipu AI

    1 crawler

AI answers

19

Fetched while an assistant was answering someone's question.

Training

20

Collected as corpus that may be used to train a model.

What this directory leaves out, on purpose.

Two robots.txt tokens

Google-Extended and Applebot-Extended control whether already-crawled content may train a model. Neither ever makes a request, so neither can appear in your log. Listing them would put a visit on your dashboard that did not happen.

Crawlers nobody has seen

Where an operator publishes no crawler documentation, we match only the tokens that are actually attested. A bare token like Grok would fire on any user agent that merely contains the word, and report a browser extension as AI traffic.

Link previews counted as AI

facebookexternalhit, Twitterbot and Slackbot fetch a page to draw a link card. Every share makes one, so folding them into the AI figure inflates it visibly. They are counted as what they are.

See which of these actually reach your pages.

An AI crawler takes your HTML and leaves — it never runs JavaScript, so a browser tag cannot see it. The server-side collector reports the fetch, verifies the address, and shows which pages the assistants are reading, including the ones that answered with a 404.