Deep dive

The AI crawler table

Search, training and user-initiated retrieval can use different controls. Use robots.txt and the linked provider documentation to choose access rules.

Latest documentation review · . Open source details in each row for fetch times.

Looking for platform-specific rules? See the dated comparison for Reddit, LinkedIn, Facebook, Instagram and TikTok. Published rules and successful content access are separate observations.

Separate search from training

Choose rules by purpose using the rows below. A search crawl block does not guarantee that a site can never be mentioned in an answer. The provider comparison explains the different routes.

Google Search and Google-Extended

Google-Extended controls specified Gemini training and grounding uses. Google says it does not affect Google Search inclusion or ranking. It is a control token in robots.txt, not a separate request user-agent. See the Google row below for the source.

What the categories mean

Model training
Collects content for model development. The provider defines the scope of any opt-out.
Search index
Collects content for search results. Check the provider’s rules for indexing and display.
User-initiated fetch
Retrieves content for a user request. robots.txt treatment differs between providers.
Ad verification
Checks landing pages for advertising policy compliance.
General / unspecified
Collects content for broader uses, such as research or product development.

Selected crawlers and controls

This is a selection of documented policies, not a complete list or a test of crawler compliance. Read how training and AI search differ across providers, including Grok, where the cited source does not establish a crawler token.

robots.txt token Operator Purpose Documented purpose and controls
bingbot Microsoft Search index

Bing’s main crawler. Crawling supports its search index; page-level directives can also control display and AI use. See the provider comparison for those separate controls.

GPTBot OpenAI Model training

Disallowing GPTBot signals that content should not be used for foundation-model training. The search setting is independent.

OAI-SearchBot OpenAI Search index

Opted-out sites are excluded from ChatGPT search answers, but may still appear as navigational links.

ChatGPT-User robots.txt may not apply OpenAI User-initiated fetch

Fetches pages for user actions. OpenAI says robots.txt may not apply; this agent does not control Search eligibility.

OAI-AdsBot OpenAI Ad verification

Visits submitted ad landing pages for safety checks and relevance. The collected data is not used for foundation-model training.

ClaudeBot Anthropic Model training

Restricting access signals that future material should be excluded from model-training datasets.

Claude-SearchBot Anthropic Search index

Disabling access prevents indexing for Claude search and may reduce search visibility.

Claude-User Anthropic User-initiated fetch

Anthropic documents robots.txt controls that prevent retrieval in response to user queries.

Googlebot Google Search index

Crawl preferences affect Google Search and its features. Google-Extended is a separate content-use control with a different product scope.

Google-Extended token only Google Training and grounding

Controls use for future Gemini training and specified Gemini/Vertex grounding features. It does not affect Google Search inclusion or ranking.

GoogleOther Google General / unspecified

A generic crawler for Google product teams, including research. Its crawl preferences are not tied to a specific product.

PerplexityBot Perplexity Search index

Surfaces and links sites in search. Perplexity recommends allowing it for search access; it is not a foundation-model training crawler.

Perplexity-User generally ignores robots.txt Perplexity User-initiated fetch

Retrieves pages for user questions. Perplexity says this fetcher generally ignores robots.txt; it is not a training crawler.

CCBot Common Crawl Open web dataset

Collects an open web-crawl dataset. Common Crawl documents a robots.txt opt-out from crawling.

Applebot Apple Search and other uses

Supplies Apple search features and can provide context for AI answers. Apple documents separate controls for training and answer context; allowing a search crawl alone does not describe every use.

Applebot-Extended token only Apple Model training

Controls training use of content collected by Applebot. It does not crawl pages itself or control search inclusion.

Amazonbot Amazon Product and model development

Collects content to improve Amazon products and services, including possible AI-model training. Amazon documents robots.txt controls.

Amzn-SearchBot Amazon Search index

Supports search experiences such as Alexa, not generative-model training. If no specific rule exists, Amazon says it follows rules given to other search bots.

Amzn-User robots.txt may not apply Amazon User-initiated fetch

Fetches current information for user requests, not generative-model training. Amazon says it may not follow all robots.txt directives.

Reading the table

User requests have different rules

OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says its user fetcher generally ignores it. Amazon documents exceptions for Amzn-User. Anthropic documents a robots.txt control for Claude-User. Use each row’s source when choosing a policy.

A control token may not appear in requests

Google-Extended governs use of content collected by existing Google crawlers. Applebot-Extended governs training use of Applebot’s data. Neither is a separate crawler to count in access logs.

Confirm requests as well as settings

When measuring activity, use the provider’s published verification method, such as IP ranges or DNS checks. A user-agent name alone is not proof of origin. Common Crawl’s linked documentation explicitly warns about impersonators.

Check firewall and CDN logs for blocked or challenged requests from crawlers you intend to allow, then confirm that a permitted request receives the useful page content. For example, Perplexity’s firewall guidance combines user-agent matching with its published IP ranges. Follow each provider’s verification method when adjusting rules.

Sources are linked in each row. Crawler behaviour changes; consult the operator’s current documentation before acting. See the practical access checklist to check a page in Google and Bing.

For the date labels beside each source, see how we date our sources.