Key takeaways
- AI crawlers split into three purposes (training, AI search indexing, and live user-triggered fetches), and blocking them is three separate decisions, not one
- Blocking training bots such as GPTBot and Google-Extended keeps your content out of model training; blocking search bots such as OAI-SearchBot and PerplexityBot removes you from AI answers and the referral traffic that comes with them
- Google-Extended does not affect Googlebot, Google Search crawling, or your rankings: it is a training opt-out only
- robots.txt is a request that reputable crawlers honour, not enforcement, and a bot-specific User-agent group overrides the wildcard group entirely
"Should I block AI crawlers in robots.txt?" is really three decisions, not one. AI crawlers do three different jobs (model training, AI search, and live user-triggered fetches), and blocking each one costs you something different. This guide shows how the bots divide up, what every rule really costs, and how to write robots.txt so the outcome is the one you intended.
The three purposes, and why they matter
Every AI crawler you will meet is doing one of three jobs. Knowing which one is the whole game.
| Purpose | What the bot does | What blocking it costs you |
|---|---|---|
| Training | Collects pages to train a foundation model | Your content is excluded from that model's training data. No effect on traffic. |
| AI search | Indexes pages so an assistant can surface and cite them | You disappear from that engine's answers and lose the referral clicks that come with a citation. |
| User-triggered | Fetches one page live because a user asked about it | A user who pastes your URL into the assistant gets an error instead of your page. |
The asymmetry is the point. Blocking training bots costs you nothing in traffic. Blocking search bots costs you the traffic. They are frequently conflated, and a robots.txt written to "keep our content out of AI" often quietly removes the site from AI answers as well.
Here is how the major tokens sort:
| Token | Operator | Purpose |
|---|---|---|
GPTBot |
OpenAI | Training |
OAI-SearchBot |
OpenAI | AI search |
ChatGPT-User |
OpenAI | User-triggered |
ClaudeBot |
Anthropic | Training |
Claude-SearchBot |
Anthropic | AI search |
Claude-User |
Anthropic | User-triggered |
Google-Extended |
Training | |
PerplexityBot |
Perplexity | AI search |
Perplexity-User |
Perplexity | User-triggered |
Applebot-Extended |
Apple | Training |
Meta-ExternalAgent |
Meta | Training |
Amazonbot |
Amazon | Training |
Bytespider |
ByteDance | Training |
CCBot |
Common Crawl | Open dataset |
Tokens and behaviour change; each operator publishes its own list, and OpenAI's, Anthropic's, Google's and Perplexity's are the ones worth re-reading occasionally. This table was checked in August 2026.
Heads up
Google-Extended is the single most misunderstood token in this list. It controls only whether your content trains Gemini and Vertex AI. It is not read by Googlebot, it does not affect crawling for Google Search, and disallowing it has no effect on your rankings. Blocking it is a data decision, not an SEO one.
How the rules are actually resolved
Two mechanics catch people out, and both silently produce the opposite of what was intended.
A bot-specific group overrides the wildcard group completely. Crawlers do not merge rules. If a bot finds a User-agent group naming it, that group is the only one it obeys. The User-agent: * group is ignored entirely for that bot. So this:
User-agent: *
Disallow: /private/
User-agent: GPTBot
Disallow: /drafts/
does not keep GPTBot out of /private/. GPTBot matches its own group, obeys only Disallow: /drafts/, and crawls /private/ freely. If a rule should apply to a named bot, it has to be repeated in that bot's group.
Disallow: with an empty value means "allow everything". It is not a no-op and it is not a comment. Disallow: / blocks the whole site; Disallow: on its own opens it.
Tip
After any robots.txt edit, re-read it as though you were the bot: find the group that names you, and read only that group. Almost every robots.txt mistake becomes obvious the moment you resolve it the way a crawler does instead of the way it looks on the page.
Three configurations that make sense
Pick the one matching your actual goal, rather than assembling rules bot by bot.
1. Open to everything. The default, and the right answer for most marketing sites and publishers who want citations. No AI rules at all, or an explicit empty disallow to make the intent legible to whoever reads the file next.
2. Out of training, into answers. The most common considered position: you would rather not fund model training with your archive, but you very much want to be cited and clicked.
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: Amazonbot
Disallow: /
Note what is absent: OAI-SearchBot, Claude-SearchBot and PerplexityBot are all still allowed, so your pages can still be surfaced and linked in AI answers.
3. Out of everything. For sites where the content is the product: paywalled journalism, proprietary research, member-only reference material. Add the search and user tokens to the list above. Understand the trade: you will not appear in AI answers, and a user who pastes your URL into an assistant will get an error.
Whichever you choose, keep your Sitemap: directive in the file. It is unaffected by any of this. Our free robots.txt generator builds a valid file from whichever configuration you pick, rather than you hand-typing user-agent blocks.
Heads up
robots.txt is a stated preference that reputable crawlers honour. Every operator above documents the token it respects. It is not enforcement, though. A crawler that ignores robots.txt is not stopped by it, and if you need a hard block the mechanism is server- or CDN-level, by user agent and verified IP range.
Checking what your rules actually do
Reading your own robots.txt is unreliable, for the reason above: the file as written and the file as resolved by a particular bot are different documents. The only trustworthy check is per-bot.
Crawl Cove's AI crawler access checker is a free browser tool that reads any site's live robots.txt and gives you a verdict for each bot in the table above, along with the exact rule that produced it. It is the quickest way to confirm that a rule you wrote for GPTBot is not being quietly overridden, and that a search bot you meant to allow has not been swept up by a wildcard. For the full, maintained list of tokens with what each one is for, see the AI crawler user-agents list. If you want the access verdict for one page together with the other things an answer engine checks (noindex and nosnippet rules, structure, author and date markup, /llms.txt), the free AI search visibility checker does that in one pass.
For a whole site, Crawl Cove's AI search visibility checker reports the same access question as findings rather than as a lookup: it flags when GPTBot, ClaudeBot, PerplexityBot or Google-Extended are disallowed from your site root, so an unintended block shows up next to your other technical issues instead of waiting to be noticed. It also raises a low-severity advisory when there is no /llms.txt file. The crawl runs locally on your own machine, so you can re-run it after every robots.txt edit and compare audits to confirm the change did what you meant. Try it free for 14 days.
Related reading
- How to write a robots.txt file covers the general syntax, rule resolution and common mistakes this post assumes.
- What is llms.txt? covers the curation side of the same question.
- Technical SEO: the complete guide shows where robots.txt sits among the rest.
- Crawl budget explained looks at how crawler access and crawl economics interact.
Wrap-up
AI crawler policy is three decisions wearing one coat. Training bots cost you nothing in traffic when you block them; AI search bots cost you citations and clicks; user-triggered bots cost you the visitor who asked about you by name. Decide those separately, remember that a bot-specific group replaces the wildcard group rather than adding to it, and then verify the file per bot rather than by eye. The rules are easy. Confirming they resolve the way you meant is the part worth the ten minutes.
Frequently asked questions
- How do I stop ChatGPT from using my website?
- Disallow GPTBot to opt out of OpenAI model training. To also keep your pages out of ChatGPT's search results and live browsing, disallow OAI-SearchBot and ChatGPT-User as well. They are separate tokens with separate purposes.
- Does blocking Google-Extended hurt my Google rankings?
- No. Google-Extended controls only whether your content trains Google's Gemini and Vertex AI models. Googlebot crawls and ranks your site independently, and disallowing Google-Extended has no effect on Search.
- Does robots.txt actually stop AI crawlers?
- Reputable operators honour it, and all the major ones publish the token they respect. But robots.txt is a stated preference, not a technical control, so it cannot prevent a crawler that chooses to ignore it. For that you need server-level blocking.
- Should I block CCBot?
- It depends on what you are optimising for. CCBot builds the open Common Crawl dataset, which is reused by a very large number of downstream models and researchers. Blocking it removes you from that corpus broadly, which is a bigger and less reversible step than blocking one company's crawler.