Skip to content

Reference

AI Crawler User-Agents List

The exact user-agent token for every major AI crawler, who runs it, and what it is actually for: training a model, powering an AI search index, or fetching a page live for a chatbot user. Updated when a real change happens, not on a fixed schedule.

OpenAI

GPTBot AI training

Crawls the web to train OpenAI foundation models (ChatGPT).

OpenAI's own documentation →
OAI-SearchBot AI search

Indexes pages to surface and link them in ChatGPT search. Not used for training.

OpenAI's own documentation →
ChatGPT-User User-triggered

Fetches a page live when a ChatGPT user asks about it. Not used for training.

OpenAI's own documentation →
OAI-AdsBot Ad validation

Validates the safety of web pages submitted as ads on ChatGPT. Does not train models or affect search.

OpenAI's own documentation →

Anthropic

ClaudeBot AI training

Crawls the web to help train Claude models.

Anthropic's own documentation →
Claude-SearchBot AI search

Indexes pages so Claude can cite them in answers.

Anthropic's own documentation →
Claude-User User-triggered

Fetches a page live in response to a Claude user request.

Anthropic's own documentation →
anthropic-ai AI training

Legacy Anthropic training token, still worth disallowing for older configs.

Anthropic's own documentation →

Google

Google-Extended AI training

Controls whether your content trains Gemini / Vertex AI. Does NOT affect Google Search crawling or ranking.

Google's own documentation →
Googlebot Search engine

The classic Google Search crawler, shown for reference so you can see AI vs. Search access side by side.

Google's own documentation →

Perplexity

PerplexityBot AI search

Indexes pages to cite them in Perplexity answers.

Perplexity's own documentation →
Perplexity-User User-triggered

Fetches a page live when a Perplexity user requests it.

Perplexity's own documentation →

Apple

Applebot-Extended AI training

Controls whether Applebot-crawled content trains Apple Intelligence models. Does not affect Siri/Spotlight lookups.

Apple's own documentation →

Meta

Meta-ExternalAgent AI training

Crawls the web to train Meta AI / Llama models.

Meta's own documentation →
Meta-ExternalFetcher User-triggered

Fetches an individual link live when a Meta AI user asks about it. Not used for training.

Meta's own documentation →

Amazon

Amazonbot AI training

Crawls the web for Amazon services including AI/Alexa answers.

Amazon's own documentation →

ByteDance (TikTok)

Bytespider AI training

Aggressive crawler collecting data for ByteDance AI models.

Common Crawl

CCBot AI dataset

Builds the open Common Crawl dataset that trains many LLMs. Blocking it removes you from a widely-reused corpus.

Common Crawl's own documentation →

DuckDuckGo

DuckAssistBot AI search

Fetches pages for DuckDuckGo’s AI-assisted answers.

DuckDuckGo's own documentation →

Mistral AI

MistralAI-Training AI training

Crawls the web to build datasets for training Mistral’s generative AI models.

Mistral AI's own documentation →
MistralAI-Index AI search

Indexes content for Mistral’s search feature in Le Chat. Not used for model training.

Mistral AI's own documentation →
MistralAI-User User-triggered

Fetches a page live when a Le Chat user asks about it. Not used for training or indexing.

Mistral AI's own documentation →

You.com

YouBot AI search

Crawls and indexes pages for You.com’s AI-powered search engine and summaries.

You.com's own documentation →

How to read this list

AI training and AI dataset bots crawl the web to build or fine-tune a model; blocking them stops your content being used for training but does nothing to your search visibility. AI search bots index pages so an AI product can cite or link to them, closer in effect to a traditional search engine crawler. User-triggered bots only fetch a page when a real person asks their AI assistant about it there and then.

robots.txt is a request, not a wall. Well-behaved crawlers honour it; nothing forces them to. And two of the entries above, Google-Extended and Applebot-Extended, only control AI training. Blocking either has no effect on Google Search or Siri/Spotlight crawling and ranking, which use separate, unrelated crawlers.

Want to know whether your own site's robots.txt currently allows or blocks each of these, and whether you have an llms.txt file at all? Run it through the free AI Crawler Access Checker.

Frequently asked questions

How often is this list updated?
When a real change happens, such as a new crawler launching or an operator renaming a token, not on a fixed schedule. Every entry here is also what our own AI Crawler Access Checker tests against, so the two never drift apart.
Does blocking Google-Extended hurt my Google rankings?
No. Google-Extended controls only whether your content can be used to train Gemini and Vertex AI models. It does not affect Googlebot, and it has no effect on Google Search crawling, indexing or ranking.
Should I block every AI crawler?
That depends on what you want. Blocking training and dataset bots (GPTBot, ClaudeBot, CCBot, Bytespider and similar) stops your content training outside models without touching search visibility. Blocking AI search bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot) removes you from citations in that assistant's answers, similar to blocking a search engine. Blocking user-triggered bots (ChatGPT-User, Claude-User, Perplexity-User) can break a page loading correctly when a real person asks their assistant about it live.
Does having an llms.txt file help me rank or appear in AI answers?
Not currently. No major AI provider has committed to reading llms.txt, and it has no effect on Google's AI Overviews or AI Mode. It is worth having as a tidy, honest summary of your site, not as a ranking lever.

Stop guessing.
Start fixing.

Crawl Cove runs on your machine, connects to your real ranking data, and tells you exactly what to fix first. No per-feature paywalls, no spreadsheets, no guesswork.

28 days risk-free · No card required