OpenAI

GPTBot

TrainingIP ranges + reverse DNS

Collects public pages that may improve future OpenAI models.

When you see it: OpenAI is gathering training data. This is the one to block if you object to training but want to stay in ChatGPT search.

How we check it is really OpenAI

  1. 01

    The user agent has to contain GPTBot.

  2. 02

    The address is compared against OpenAI's published ranges, which we refetch every six hours.

  3. 03

    If the ranges do not settle it, the address's reverse DNS record is resolved and then resolved forward again, so a faked record cannot pass.

A visit that matches neither is recorded as unverified, not as a forgery. A range list can be incomplete, and an accusation needs evidence.

What each result means

verified
The address matched OpenAI's own published records. This really was them.
unverified
We could not confirm it, and that is all it means. The operator may publish nothing to check against, a published list may be behind the addresses actually in use, or DNS may simply have been slow. It is not an accusation.
spoofed
The address belongs to somebody else — its own reverse DNS record resolves to a different owner. That is positive evidence of a forgery rather than a failure to find evidence, which is why it is reported as its own result and never folded into the row above.

The distinction costs nothing to make and is the reason these numbers can be shown to someone who will ask where they came from. A report that cannot separate “we do not know” from “this was faked” invites exactly one question it cannot answer.

The lists this is checked against

Published by OpenAI and refetched every six hours. Open them; they are the same files we read.

Operator
OpenAI
Token
GPTBot
Counted as
AI training
Verification
IP ranges + reverse DNS
Runs JavaScript
No
Operator docs
Published

User agent

GPTBot

Matched as a token, not as a full browser string. OpenAI adds and changes the version and product text around it without notice, so a filter pinned to the whole header stops working the first time they bump a version while one matching GPTBot keeps working. The same string is what OpenAI names in robots.txt.

One real request

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot

OpenAI's other crawlers

OpenAI sends 5 crawlers, each with its own token and its own job. They are counted separately for that reason — the tokens differ by a few characters and what the fetch means does not, so folding them into one operator row would let indexing look like AI interest.

TokenCounted asVerificationWhat it fetches for
GPTBotthis pageTrainingIP ranges + reverse DNSCollects public pages that may improve future OpenAI models.
ChatGPT-UserAI answersIP ranges + reverse DNSFetches a page while ChatGPT is answering someone's question.
ChatGPT-AgentAI answersIP ranges + reverse DNSChatGPT's agent mode browsing on a person's behalf.
OAI-SearchBotIndexingIP ranges + reverse DNSBuilds the index behind ChatGPT search.
OAI-AdsBotOther AI trafficIP ranges + reverse DNSChecks pages submitted as ChatGPT ads.

Your analytics never mentioned GPTBot because it cannot see it.

This crawler takes your HTML and leaves without running a line of JavaScript, so a browser tag records nothing. TrueStat reads it server-side, checks the address, and shows you which pages it fetched — including the ones that returned a 404.