| bingbot | Microsoft | Search index | Bing’s main crawler. Crawling supports its search index; page-level directives can also control display and AI use. See the provider comparison for those separate controls. |
| GPTBot | OpenAI | Model training | Disallowing GPTBot signals that content should not be used for foundation-model training. The search setting is independent. |
| OAI-SearchBot | OpenAI | Search index | Opted-out sites are excluded from ChatGPT search answers, but may still appear as navigational links. |
| ChatGPT-User robots.txt may not apply | OpenAI | User-initiated fetch | Fetches pages for user actions. OpenAI says robots.txt may not apply; this agent does not control Search eligibility. |
| OAI-AdsBot | OpenAI | Ad verification | Visits submitted ad landing pages for safety checks and relevance. The collected data is not used for foundation-model training. |
| ClaudeBot | Anthropic | Model training | Restricting access signals that future material should be excluded from model-training datasets. |
| Claude-SearchBot | Anthropic | Search index | Disabling access prevents indexing for Claude search and may reduce search visibility. |
| Claude-User | Anthropic | User-initiated fetch | Anthropic documents robots.txt controls that prevent retrieval in response to user queries. |
| Googlebot | Google | Search index | Crawl preferences affect Google Search and its features. Google-Extended is a separate content-use control with a different product scope. |
| Google-Extended token only | Google | Training and grounding | Controls use for future Gemini training and specified Gemini/Vertex grounding features. It does not affect Google Search inclusion or ranking. |
| GoogleOther | Google | General / unspecified | A generic crawler for Google product teams, including research. Its crawl preferences are not tied to a specific product. |
| PerplexityBot | Perplexity | Search index | Surfaces and links sites in search. Perplexity recommends allowing it for search access; it is not a foundation-model training crawler. |
| Perplexity-User generally ignores robots.txt | Perplexity | User-initiated fetch | Retrieves pages for user questions. Perplexity says this fetcher generally ignores robots.txt; it is not a training crawler. |
| CCBot | Common Crawl | Open web dataset | Collects an open web-crawl dataset. Common Crawl documents a robots.txt opt-out from crawling. |
| Applebot | Apple | Search and other uses | Supplies Apple search features and can provide context for AI answers. Apple documents separate controls for training and answer context; allowing a search crawl alone does not describe every use. |
| Applebot-Extended token only | Apple | Model training | Controls training use of content collected by Applebot. It does not crawl pages itself or control search inclusion. |
| Amazonbot | Amazon | Product and model development | Collects content to improve Amazon products and services, including possible AI-model training. Amazon documents robots.txt controls. |
| Amzn-SearchBot | Amazon | Search index | Supports search experiences such as Alexa, not generative-model training. If no specific rule exists, Amazon says it follows rules given to other search bots. |
| Amzn-User robots.txt may not apply | Amazon | User-initiated fetch | Fetches current information for user requests, not generative-model training. Amazon says it may not follow all robots.txt directives. |