Which robots.txt line blocks your AI crawlers?
By Grady Coleman (Founder) · Last updated:
Three rules decide everything: the most specific user-agent group applies, the longest matching path rule wins, and Allow beats Disallow at equal length. Here is how to read your file like the checker does — and fix only what needs fixing.
Check your domain nowKey takeaways
- — The most specific user-agent group applies; longest matching path wins.
- — Training bots and search bots are independent controls — decide each deliberately.
- — A 403 means refused by some layer; a 429 means slowed down, not blocked.
- — Verify the live file after every change: deploys and apps overwrite it.
Which group applies to each bot?
A group starting with User-agent: GPTBot applies only to GPTBot. A group with User-agent: * applies to any bot with no specific group. The token must match exactly — GPT-Bot, ChatGPT, or OpenAI match nothing. Our checker evaluates all 9 canonical identities separately so a training-bot block is never confused with a search-bot block.
Should you block training, search, or neither?
This is the decision most sites get wrong — they block everything when they only meant to opt out of training:
Allow for AI search visibility: OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, Bingbot.
Decide deliberately (training): GPTBot, ClaudeBot, Google-Extended.
Never block blindly: live-fetch agents that act for a real user at request time.
Copy-paste rules for each AI crawler
The token must match the crawler's published user-agent exactly, and an empty Disallow: means allow — that is the whole trick. Paste into your robots.txt, save, then open the live file at yourdomain.com/robots.txt to confirm the change survived, because CMS apps and deploys can overwrite it.
Block GPTBot — opt out of OpenAI model training
User-agent: GPTBot Disallow: /
Only the exact token GPTBot is affected. OAI-SearchBot and every other crawler are separate groups and keep working. For the full GPTBot profile — what it does, and what blocking does not do — see the GPTBot guide.
Allow GPTBot explicitly — after a wildcard block
User-agent: GPTBot Disallow:
A group for the exact token overrides the User-agent: * group for that crawler, so this re-opens GPTBot even when * is disallowed.
Allow the search and retrieval crawlers
User-agent: OAI-SearchBot Disallow: User-agent: PerplexityBot Disallow: User-agent: Claude-SearchBot Disallow:
These three power ChatGPT search, Perplexity retrieval, and Claude web search. Blocking one does not block the others, and access never guarantees appearance in results or answers.
Block ClaudeBot — opt out of Anthropic training
User-agent: ClaudeBot Disallow: /
ClaudeBot and Claude-SearchBot are independent controls. Blocking training does not remove you from Claude search if the search bot is allowed.
Block Google-Extended — opt out of Gemini training use
User-agent: Google-Extended Disallow: /
Google documents this as a control token for Gemini training use, not a separate crawling bot — it does not affect Google Search ranking.
The full recipe: opt out of training, stay visible in AI search
# Search and retrieval: allowed User-agent: OAI-SearchBot Disallow: User-agent: PerplexityBot Disallow: User-agent: Claude-SearchBot Disallow: # Training: blocked User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: /
Each group is evaluated per token, so this keeps the search crawlers working while opting out of the three training controls. Verify the live file afterwards.
OAI-AdsBot — only if you run ChatGPT Ads
User-agent: OAI-AdsBot Disallow:
This token validates ChatGPT Ads landing pages. If you do not advertise there, its state is irrelevant to organic visibility.
Perplexity also documents Perplexity-User, a user-triggered fetcher that acts at request time when someone asks Perplexity about your page — treat it like a live-fetch agent rather than a scheduled crawler (see their crawler documentation below). And if you need the identity of each token, the crawler reference lists all nine.
What do 403 and 429 actually mean?
A 403 observed at fetch time was refused by the server or firewall — it does not name Cloudflare, Vercel, or any layer by itself. A 429 means rate limiting was observed, not a permanent block. And a homepage result never establishes the whole domain: re-check the exact page path the client cares about.
Limits
Point-in-time configuration evidence for the submitted page path only. Not proof of a crawler visit, firewall behavior, indexing, training, ranking, citation, traffic, or revenue. No tool can promise an AI answer will cite you.
Frequently asked questions
How do I block GPTBot in robots.txt?
Add a group that names the token exactly: User-agent: GPTBot followed by Disallow: / blocks the whole site for GPTBot. Only that exact token is affected — OAI-SearchBot, PerplexityBot, and every other crawler are separate groups and keep working. After saving, open yourdomain.com/robots.txt in a browser and confirm the lines are live, because CMS apps and deploys can overwrite the file.
What is the difference between GPTBot and OAI-SearchBot?
According to OpenAI's crawler documentation, GPTBot crawls content that may be used for model training, while OAI-SearchBot crawls specifically for ChatGPT search results. They are independent robots.txt user-agents: you can block training crawls with User-agent: GPTBot / Disallow: / while allowing search crawls with User-agent: OAI-SearchBot and an empty Disallow. Blocking one does not block the other.
How do I allow an AI crawler in robots.txt?
The fix is a group for the crawler's exact user-agent token with an empty Disallow line, e.g. User-agent: PerplexityBot followed by Disallow: (empty). An empty Disallow means allowed — that is not a typo. Verify the live file at yourdomain.com/robots.txt afterwards, because CMS apps and deploys can overwrite it.
What does a 403 mean for an AI crawler?
The 403 status observed at fetch time means access was refused by the server or firewall layer — it does not name the layer. robots.txt itself never returns 403; it is a text file. Say '403 observed', not 'Cloudflare blocked the crawler', unless you have rule-level evidence from that layer.
What does a 429 mean?
The 429 response means rate limiting was observed — the crawler asked too fast and was slowed down. It is not permanent blocking. Recheck off-peak before concluding anything, and do not treat one 429 as proof the bot can never reach you.
Does blocking AI training bots remove me from AI answers?
The bots you block are the only ones that stop being asked. Training bots (GPTBot, ClaudeBot) and search bots (OAI-SearchBot, PerplexityBot, Claude-SearchBot) are separate controls. If you block only training, search-index bots you allowed can still read you. Check each token individually rather than blocking everything at once.
By Grady Coleman (Founder) · Last updated: