Reference
AI Crawler User-Agents List
The exact user-agent token for every major AI crawler, who runs it, and what it is actually for: training a model, powering an AI search index, or fetching a page live for a chatbot user. Updated when a real change happens, not on a fixed schedule.
OpenAI
GPTBot
AI training
Crawls the web to train OpenAI foundation models (ChatGPT).
OpenAI's own documentation →OAI-SearchBot
AI search
Indexes pages to surface and link them in ChatGPT search. Not used for training.
OpenAI's own documentation →ChatGPT-User
User-triggered
Fetches a page live when a ChatGPT user asks about it. Not used for training.
OpenAI's own documentation →OAI-AdsBot
Ad validation
Validates the safety of web pages submitted as ads on ChatGPT. Does not train models or affect search.
OpenAI's own documentation →Anthropic
Claude-SearchBot
AI search
Indexes pages so Claude can cite them in answers.
Anthropic's own documentation →Claude-User
User-triggered
Fetches a page live in response to a Claude user request.
Anthropic's own documentation →anthropic-ai
AI training
Legacy Anthropic training token, still worth disallowing for older configs.
Anthropic's own documentation →Google-Extended
AI training
Controls whether your content trains Gemini / Vertex AI. Does NOT affect Google Search crawling or ranking.
Google's own documentation →Googlebot
Search engine
The classic Google Search crawler, shown for reference so you can see AI vs. Search access side by side.
Google's own documentation →Perplexity
PerplexityBot
AI search
Indexes pages to cite them in Perplexity answers.
Perplexity's own documentation →Perplexity-User
User-triggered
Fetches a page live when a Perplexity user requests it.
Perplexity's own documentation →Apple
Applebot-Extended
AI training
Controls whether Applebot-crawled content trains Apple Intelligence models. Does not affect Siri/Spotlight lookups.
Apple's own documentation →Meta
Meta-ExternalAgent
AI training
Crawls the web to train Meta AI / Llama models.
Meta's own documentation →Meta-ExternalFetcher
User-triggered
Fetches an individual link live when a Meta AI user asks about it. Not used for training.
Meta's own documentation →Amazon
Amazonbot
AI training
Crawls the web for Amazon services including AI/Alexa answers.
Amazon's own documentation →ByteDance (TikTok)
Bytespider
AI training
Aggressive crawler collecting data for ByteDance AI models.
Common Crawl
CCBot
AI dataset
Builds the open Common Crawl dataset that trains many LLMs. Blocking it removes you from a widely-reused corpus.
Common Crawl's own documentation →DuckDuckGo
DuckAssistBot
AI search
Fetches pages for DuckDuckGo’s AI-assisted answers.
DuckDuckGo's own documentation →Mistral AI
MistralAI-Training
AI training
Crawls the web to build datasets for training Mistral’s generative AI models.
Mistral AI's own documentation →MistralAI-Index
AI search
Indexes content for Mistral’s search feature in Le Chat. Not used for model training.
Mistral AI's own documentation →MistralAI-User
User-triggered
Fetches a page live when a Le Chat user asks about it. Not used for training or indexing.
Mistral AI's own documentation →You.com
YouBot
AI search
Crawls and indexes pages for You.com’s AI-powered search engine and summaries.
You.com's own documentation →How to read this list
AI training and AI dataset bots crawl the web to build or fine-tune a model; blocking them stops your content being used for training but does nothing to your search visibility. AI search bots index pages so an AI product can cite or link to them, closer in effect to a traditional search engine crawler. User-triggered bots only fetch a page when a real person asks their AI assistant about it there and then.
robots.txt is a request, not a wall. Well-behaved crawlers
honour it; nothing forces them to. And two of the entries above,
Google-Extended
and
Applebot-Extended,
only control AI training. Blocking either has no effect on Google Search or Siri/Spotlight
crawling and ranking, which use separate, unrelated crawlers.
Want to know whether your own site's robots.txt
currently allows or blocks each of these, and whether you have an
llms.txt
file at all? Run it through the free
AI Crawler Access Checker.
Frequently asked questions
- How often is this list updated?
- When a real change happens, such as a new crawler launching or an operator renaming a token, not on a fixed schedule. Every entry here is also what our own AI Crawler Access Checker tests against, so the two never drift apart.
- Does blocking Google-Extended hurt my Google rankings?
- No. Google-Extended controls only whether your content can be used to train Gemini and Vertex AI models. It does not affect Googlebot, and it has no effect on Google Search crawling, indexing or ranking.
- Should I block every AI crawler?
- That depends on what you want. Blocking training and dataset bots (GPTBot, ClaudeBot, CCBot, Bytespider and similar) stops your content training outside models without touching search visibility. Blocking AI search bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot) removes you from citations in that assistant's answers, similar to blocking a search engine. Blocking user-triggered bots (ChatGPT-User, Claude-User, Perplexity-User) can break a page loading correctly when a real person asks their assistant about it live.
- Does having an llms.txt file help me rank or appear in AI answers?
- Not currently. No major AI provider has committed to reading llms.txt, and it has no effect on Google's AI Overviews or AI Mode. It is worth having as a tidy, honest summary of your site, not as a ranking lever.
Stop guessing.
Start fixing.
Crawl Cove runs on your machine, connects to your real ranking data, and tells you exactly what to fix first. No per-feature paywalls, no spreadsheets, no guesswork.
28 days risk-free · No card required