Nine tokens. Three purposes. Know which is which.
By Grady Coleman (Founder) · Last updated:
AI robots.txt tokens are not interchangeable. A training crawler and a search crawler control different things — this reference keeps them separated so you never block search when you only meant to opt out of training.
Check all 9 on your domainKey takeaways
- — Five tokens serve search and discovery; three serve training; one validates ads.
- — Tokens with similar names do different jobs: GPTBot trains, OAI-SearchBot retrieves.
- — Google-Extended is a control token, not a separate bot, and never affects ranking.
- — Blocking training never blocks search — decide each purpose on its own.
Which crawlers handle search and discovery?
The five search tokens contribute to search and AI search experiences: Googlebot, OAI-SearchBot, PerplexityBot, Claude-SearchBot, and Bingbot. Blocking any of them may reduce the evidence available to that provider, but access alone never guarantees appearance in results or answers. According to each provider's crawler documentation, retrieval permission and result inclusion are separate decisions.
Googlebot
Google Search and AI Overviews eligibility. Required for both.
OAI-SearchBot
OpenAI
Used for OpenAI search retrieval. Access does not guarantee appearance in ChatGPT search results.
PerplexityBot
Perplexity
Used for Perplexity retrieval. Access does not guarantee inclusion in answers or citations.
Claude-SearchBot
Anthropic
Powers Claude web search. Blocking it can reduce visibility in Claude searches.
Bingbot
Microsoft
Bing index and Copilot grounding.
Which crawlers train models?
The three training tokens are associated with training data or other AI product uses: GPTBot, Google-Extended, and ClaudeBot. Blocking these does not, by itself, block search discovery — the search tokens above are independent controls. The distinction matters because most sites mean to opt out of training while staying visible in search, and these separate tokens make that possible.
GPTBot
OpenAI
Crawls content that may be used for model training. Separate from OAI-SearchBot.
Google-Extended
A control token (not a separate bot) for Gemini training use. Does not affect search ranking.
ClaudeBot
Anthropic
Crawls content for Claude training and retrieval. Separate from Claude-SearchBot.
Which crawler validates ad landing pages?
The single ad token fetches pages to validate ad landing pages: OAI-AdsBot for ChatGPT Ads. It is not an organic visibility signal, and its result says nothing about rankings, citations, or whether an ad will be approved. The token matters only when ChatGPT Ads landing pages are in play.
OAI-AdsBot
OpenAI
Validates ChatGPT Ads landing pages. Relevant only if you run ChatGPT Ads.
How do you allow or block a single token?
One group per decision. According to the robots exclusion conventions, a group beginning with the exact token (for example User-agent: GPTBot) followed by Disallow: / blocks only that crawler, while an empty Disallow allows it. The groups are independent: allowing OAI-SearchBot never allows GPTBot, and blocking ClaudeBot never blocks Claude-SearchBot. According to each provider's crawler documentation, the token must match exactly — close variants match nothing. Verify the live file after every change.
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Disallow:
How do training, search, and ad tokens compare?
The three purposes side by side. According to each provider's crawler documentation, the token — not the vendor — decides what a rule controls.
| Purpose | Tokens | Blocking it affects |
|---|---|---|
| Search and discovery | Googlebot, OAI-SearchBot, PerplexityBot, Claude-SearchBot, Bingbot | Retrieval evidence for search and AI answers |
| Training and AI use | GPTBot, Google-Extended, ClaudeBot | Model training eligibility only — never rankings |
| Ad validation | OAI-AdsBot | ChatGPT Ads landing checks only |
Which token questions come up most?
The three confusions this reference exists to prevent, answered directly.
Can I allow OAI-SearchBot while blocking GPTBot?
Yes, the two tokens are independent controls in robots.txt. According to OpenAI's crawler documentation, OAI-SearchBot serves search retrieval while GPTBot serves training crawls, so a group allowing one and disallowing the other does exactly what it says. The checker reports both states separately so the split stays visible on every scan.
Does blocking GPTBot remove me from ChatGPT search?
No, blocking GPTBot alone does not remove a site from ChatGPT search results. The search path runs through OAI-SearchBot, a separate token with separate directives. According to the token split above, training opt-outs and search visibility are decided independently — verify each token rather than assuming one rule covers both jobs.
What about live-fetch agents like ChatGPT-User?
The 9-token reference covers index crawlers and validation tokens, not per-request live fetch. Live-fetch agents retrieve a page at request time for a real user rather than crawling an index, so they follow their own behavior on top of robots directives. The distinction matters because an Allowed index result says nothing about what a live fetch will do in the moment.
Honesty box
Permission is not proof of a visit, indexing, ingestion, ranking, or citation. robots.txt is advisory — a firewall, CDN rule, or bot-management layer can silently block a crawler that robots.txt allows. Broader utilities may test additional bots as a separate experiment; the paid Optimus scope stays at these 9 identities.
By Grady Coleman (Founder) · Last updated: