Crawler directory
Every crawler we can name, and how we know it is really them.
A user agent is a claim, not proof — anyone can send curl -A "GPTBot". Each of these is checked against the operator's own published address ranges or its reverse DNS record, and where neither exists the crawler's page says so instead of implying a check we cannot perform.
- Crawlers
- 59
- Operators
- 23
- Verifiable
- 35
- Buckets
- 4
Plus 1,500 generic patterns for scrapers, monitors and headless browsers — counted in your dashboard, not named here.
23 operators, 13 of them publishing something we can check.
An operator earns a name here by being classified in the detector, not by being famous. The ones that publish nothing still get a row — they simply cannot get a tick.
Google
7 crawlers · verifiable
OpenAI
5 crawlers · verifiable
xAI
5 crawlers
Anthropic
4 crawlers · verifiable
Meta
4 crawlers
Alibaba
3 crawlers
Amazon
3 crawlers
ByteDance
3 crawlers
Microsoft Bing
3 crawlers · verifiable
Mistral
3 crawlers · verifiable
Moonshot AI
3 crawlers · verifiable
Baidu
2 crawlers · verifiable
DuckDuckGo
2 crawlers · verifiable
Huawei
2 crawlers · verifiable
Perplexity
2 crawlers · verifiable
Allen AI
1 crawler
Apple
1 crawler · verifiable
Cohere
1 crawler
Common Crawl
1 crawler · verifiable
DeepSeek
1 crawler
Yandex
1 crawler · verifiable
You.com
1 crawler
Zhipu AI
1 crawler
AI answers
19Fetched while an assistant was answering someone's question.
ChatGPT-UserFetches a page while ChatGPT is answering someone's question.IP ranges + reverse DNSChatGPT-AgentChatGPT's agent mode browsing on a person's behalf.IP ranges + reverse DNSClaude-UserFetches a page while Claude is answering someone's question.Published IP rangesClaude-CodeAnthropic's coding agent reading a page a developer pointed it at.Published IP rangesPerplexity-UserFetches a page Perplexity is about to cite in an answer.Published IP rangesGoogle-NotebookLMFetches a page a person added to a NotebookLM notebook.IP ranges + reverse DNSGoogle-AgentA Google agent fetching on a person's behalf.IP ranges + reverse DNSGemini-Deep-ResearchGemini's deep research mode gathering sources.IP ranges + reverse DNSGoogle-Read-AloudFetches a page so it can be read out loud.IP ranges + reverse DNSCopilotMicrosoft Copilot fetching a page to answer with.IP ranges + reverse DNSBingPreviewBuilds the preview Copilot shows beside a cited link.IP ranges + reverse DNSMistralAI-UserLe Chat fetching a page to answer someone's question.Published IP rangesKimi-UserKimi fetching a page to answer someone's question.Published IP rangesDuckAssistBotFetches a page DuckDuckGo is citing in an AI-assisted answer.Published IP rangesAmzn-UserAmazon fetching a page for a live answer, including Alexa.User agent onlymeta-externalfetcherFetches a specific URL a person asked Meta AI about.User agent onlyxAI-SearchBotGrok fetching a page while answering.User agent onlyGrok-DeepSearchGrok's deep research mode gathering sources for a long answer.User agent onlyQwen-UserQwen fetching a page to answer someone's question.User agent only
Indexing
15Crawled so a search or answer engine can find the page later.
GooglebotGoogle's search crawler, and the input to AI Overviews.IP ranges + reverse DNSOAI-SearchBotBuilds the index behind ChatGPT search.IP ranges + reverse DNSClaude-SearchBotDiscovers and refreshes content for Claude's search.Published IP rangesPerplexityBotKeeps Perplexity's search index fresh. It is an index, not training.User agent onlybingbotMicrosoft's search crawler for Bing.IP ranges + reverse DNSApplebotServes Siri and Spotlight Suggestions.IP ranges + reverse DNSAmzn-SearchBotMakes content eligible for Amazon search experiences.User agent onlymeta-webindexerBuilds the index behind Meta AI search.User agent onlyMistralAI-IndexBuilds Mistral's search index.Published IP rangesKimi-SearchBotBuilds Kimi's search index.Published IP rangesDuckDuckBotDuckDuckGo's ordinary search crawler.Published IP rangesYandexBotYandex's search crawler.IP ranges + reverse DNSBaiduspiderBaidu's search crawler.IP ranges + reverse DNSPetalBotHuawei's Petal Search crawler.IP ranges + reverse DNSxAI-GrokRetrieves and indexes pages so Grok can find them later.User agent only
Training
20Collected as corpus that may be used to train a model.
TikTokSpiderByteDance's crawler. No operator documentation exists for it.User agent onlyGPTBotCollects public pages that may improve future OpenAI models.IP ranges + reverse DNSClaudeBotCollects public content that may be used to improve Claude.Published IP rangesGoogleOtherGoogle product teams fetching public content, including research.IP ranges + reverse DNSGoogle-CloudVertexBotCrawls sites for owners building Vertex AI agents.IP ranges + reverse DNSmeta-externalagentMeta's training corpus crawler.User agent onlyBytespiderByteDance's training crawler.User agent onlybedrockbotFetches content for Amazon Bedrock.User agent onlyMistralAI-TrainingMistral's training corpus crawler.Published IP rangesKimiBotMoonshot's training corpus crawler.Published IP rangesCCBotBuilds the open web dataset most model builders start from.Published IP rangescohere-aiCohere's crawler.User agent onlyDeepSeekBotDeepSeek's crawler.User agent onlyChatGLM-SpiderZhipu's crawler for GLM.User agent onlyAI2BotThe Allen Institute's crawler for open AI research datasets.User agent onlyQwenBotAlibaba's training crawler for Qwen.User agent onlyERNIEBotBaidu's training crawler for ERNIE.IP ranges + reverse DNSPanguBotHuawei's training crawler for Pangu.IP ranges + reverse DNSYouBotYou.com's crawler. It both indexes and collects, with no token separating the two.User agent onlyGrokBotxAI's training crawler, signed with its own address at x.ai.User agent only
Other AI traffic
5Operator known, purpose outside the three above — ad checks, cloud fetches.
OAI-AdsBotChecks pages submitted as ChatGPT ads.IP ranges + reverse DNSmeta-externaladsMeta's crawler for advertising and business products.User agent onlyAliyunBotAlibaba Cloud infrastructure traffic.User agent onlyDoubaobotByteDance's crawler for the Doubao assistant.User agent onlyxAI-BotThe family name xAI's own robots.txt examples use.User agent only
What this directory leaves out, on purpose.
Two robots.txt tokens
Google-Extended and Applebot-Extended control whether already-crawled content may train a model. Neither ever makes a request, so neither can appear in your log. Listing them would put a visit on your dashboard that did not happen.
Crawlers nobody has seen
Where an operator publishes no crawler documentation, we match only the tokens that are actually attested. A bare token like Grok would fire on any user agent that merely contains the word, and report a browser extension as AI traffic.
Link previews counted as AI
facebookexternalhit, Twitterbot and Slackbot fetch a page to draw a link card. Every share makes one, so folding them into the AI figure inflates it visibly. They are counted as what they are.
See which of these actually reach your pages.
An AI crawler takes your HTML and leaves — it never runs JavaScript, so a browser tag cannot see it. The server-side collector reports the fetch, verifies the address, and shows which pages the assistants are reading, including the ones that answered with a 404.