AI crawler directory: block or allow, per bot

40 crawlers, each read from its operator’s own documentation: the exact user-agent token, what it does with fetched content, whether the operator says it honors robots.txt, the copy-paste block directive, and — the part most lists skip — what blocking it actually costs you. Training bots and retrieval bots are different decisions: blocking a trainer keeps your words out of future models, while blocking a retrieval bot removes you from AI answers and citations today. Check what your own file does for every bot below with the AI crawler robots.txt tester, pick a rule set by goal in the robots.txt guide, and confirm the traffic is genuine with the log verification guide.

Public research data

Verified, dated, and reusable

Verified records
40
Last source check
Next review
Schema version
2026-09-26.2

Each record uses the crawler operator's own documentation. Unclear claims stay unclear. Missing official evidence is not replaced with third-party claims.

Model training crawlers

Collect content to train future models. Blocking these does not remove you from today's AI answers.

BotOperatorHonors robots.txtDetails
AI2BotAllen Institute for AInot documentedShould you block AI2Bot?
Applebot-ExtendedAppleyes (per operator)Should you block Applebot-Extended?
ClaudeBotAnthropicyes (per operator)Should you block ClaudeBot?
GPTBot/1.4OpenAIyes (per operator)Should you block GPTBot?
MistralAI-Training/1.0Mistralyes (per operator)Should you block MistralAI-Training?

Model training and AI grounding controls

Control model training and, for some operators, whether indexed content can ground current AI answers.

BotOperatorHonors robots.txtDetails
Google-ExtendedGoogleyes (per operator)Should you block Google-Extended?
meta-externalagent/1.1Metayes (per operator)Should you block meta-externalagent?

Live retrieval and user-request fetchers

Fetch pages when a user asks an assistant something. Blocking these removes you from AI answers and citations.

BotOperatorHonors robots.txtDetails
Amzn-User/0.1Amazonpartial; some requests may bypassShould you block Amzn-User?
ChatGPT-User/1.0OpenAIpartial; some requests may bypassShould you block ChatGPT-User?
Claude-UserAnthropicyes (per operator)Should you block Claude-User?
meta-externalfetcher/1.1Metapartial; some requests may bypassShould you block meta-externalfetcher?
MistralAI-User/1.0Mistralyes (per operator)Should you block MistralAI-User?
Perplexity-User/1.0PerplexitynoShould you block Perplexity-User?

Search and AI-search indexers

Build the indexes behind classic and AI search results.

BotOperatorHonors robots.txtDetails
Amzn-SearchBot/0.1Amazonyes (per operator)Should you block Amzn-SearchBot?
ApplebotAppleyes (per operator)Should you block Applebot?
bingbotMicrosoftyes (per operator)Should you block Bingbot?
Claude-SearchBotAnthropicyes (per operator)Should you block Claude-SearchBot?
DiffbotDiffbotpartial; some requests may bypassShould you block Diffbot?
DuckAssistBot/1.2DuckDuckGoyes (per operator)Should you block DuckAssistBot?
DuckDuckBotDuckDuckGoyes (per operator)Should you block DuckDuckBot?
GooglebotGoogleyes (per operator)Should you block Googlebot?
LinerBot/1.0Lineryes (per operator)Should you block LinerBot?
meta-webindexer/1.1Metayes (per operator)Should you block Meta-WebIndexer?
MistralAI-Index/1.0Mistralyes (per operator)Should you block MistralAI-Index?
OAI-SearchBot/1.4OpenAIyes (per operator)Should you block OAI-SearchBot?
PerplexityBot/1.0Perplexityyes (per operator)Should you block PerplexityBot?
PetalBotHuaweiyes (per operator)Should you block PetalBot?
TimpibotTimpinot documentedShould you block Timpibot? (unverified)
YouBot/1.0You.comyes (per operator)Should you block YouBot?

SEO and marketing tool crawlers

Feed backlink and rank databases. Blocking them hides you from tools, not from readers.

BotOperatorHonors robots.txtDetails
AhrefsBot/7.0Ahrefsyes (per operator)Should you block AhrefsBot?
DataForSeoBotDataForSEOyes (per operator)Should you block DataForSeoBot?
MJ12bot/v1.4.8Majesticyes (per operator)Should you block MJ12bot?
SemrushBotSemrushyes (per operator)Should you block SemrushBot?

Archive and dataset crawlers

Build public datasets and archives that other AI systems train on.

BotOperatorHonors robots.txtDetails
CCBot/2.0Common Crawlyes (per operator)Should you block CCBot?

Other product crawlers

Fetch public content for product research or uses that the operator does not tie to one search or AI product.

BotOperatorHonors robots.txtDetails
Amazonbot/0.1Amazonyes (per operator)Should you block Amazonbot?
facebookexternalhit/1.1Metapartial; some requests may bypassShould you block FacebookExternalHit?
GoogleOtherGoogleyes (per operator)Should you block GoogleOther?
ImagesiftBotImageSift (Hive)yes (per operator)Should you block ImagesiftBot?
meta-externalads/1.1Metayes (per operator)Should you block Meta-ExternalAds?
OAI-AdsBot/1.0OpenAInot documentedShould you block OAI-AdsBot?

Which operators run more than one crawler?

A robots.txt group binds the crawlers whose tokens it names, so blocking one of an operator’s crawlers by name usually leaves the others running. These operators split the job across several tokens; each card notes the exceptions its operator documents.

OpenAI

OpenAI documents four crawlers: GPTBot for training, OAI-SearchBot for ChatGPT search, ChatGPT-User for fetches a user asks for, and OAI-AdsBot, which only visits pages submitted as ChatGPT ads. OpenAI says each robots.txt setting is independent, publishes a separate IP file for each bot, and may reuse one crawl for both GPTBot and OAI-SearchBot when a site allows both.

Anthropic

Anthropic documents three bots with separate tokens: ClaudeBot for training, Claude-SearchBot for search indexing, and Claude-User for fetches a user asks for. Anthropic says its bots honor robots.txt and supports the Crawl-delay extension. All three come from the addresses in one published file, so an IP match confirms Anthropic but not which bot. Anthropic warns that an IP block may not work as an opt-out, because it stops the bot from reading your robots.txt.

Google

Google splits control three ways. Googlebot crawls for Search, including AI Overviews and AI Mode. GoogleOther runs product and research crawls; Google says preferences set for it don't affect any specific product. Google-Extended fetches nothing itself; it only controls whether content Google already crawls may train or ground Gemini. Googlebot and GoogleOther share one published IP file and reverse-DNS names ending in googlebot.com.

Apple

Apple's web crawler is Applebot, and Apple gives publishers three separate controls over what it collects. (A separate iTMS agent from the same hosts fetches only registered Apple Podcasts URLs and ignores robots.txt.) A robots.txt rule for Applebot controls crawling and search appearance in Spotlight, Siri, and Safari. A rule for Applebot-Extended controls only whether that crawled content trains Apple's foundation models. The nosnippet meta tag, or isAccessibleForFree set to false for paywalled pages, keeps content out of the AI-generated answers Apple shows. Apple publishes Applebot's IP ranges and reverse-DNS names under applebot.apple.com.

Perplexity

Perplexity documents two agents that work independently: PerplexityBot indexes pages for Perplexity's search results, and Perplexity-User visits pages when a user asks a question. Perplexity says neither is used to crawl content for AI foundation models. Changes to robots.txt can take up to 24 hours to apply, and each agent has its own published IP file.

DuckDuckGo

DuckDuckGo documents two crawlers that share one address pool: DuckDuckBot for its own search indexes and DuckAssistBot for its AI-assisted answers. Each has its own robots.txt token, so you can refuse the AI answers and stay in search. Because the IPs overlap, the user-agent token is the only way to tell them apart in your logs.

Meta

Meta documents five crawlers, and each is blocked by its own robots.txt token, so a rule for one leaves the others untouched. Meta says its crawlers may cache robots.txt for up to 24 hours, so a change can take a day to apply. It publishes user-agent strings but no IP list, and it points site owners to robots.txt rather than non-standard signals such as NoAI tags.

Amazon

Amazon documents three agents with independent robots.txt settings: Amazonbot for product improvement and possible AI training, Amzn-SearchBot for search experiences such as Alexa, and Amzn-User for live fetches on a customer's behalf. Amazon says the two search and live agents do not crawl for generative AI training. Amazon says a settings change may take about 24 hours to apply, though all three may use a robots.txt copy cached for up to 30 days. None supports Crawl-delay, and Amazon publishes a separate IP list for each.

Mistral

Mistral documents three agents, each with its own robots.txt token: MistralAI-Training builds training datasets, MistralAI-Index builds the index behind Mistral search, and MistralAI-User fetches pages when a Vibe user asks. Mistral says the index and user agents are not used for training, and the training crawler serves neither search nor live answers, so each rule does one job. Mistral publishes IP files for the index and user agents but not for the training crawler.

Change history

This log records changes to the directory method and coverage. Each crawler record also keeps its own source check and review date.

  1. Published the first verified directory with official operator sources, robots.txt controls, verification methods, and scheduled review dates.
  2. Omitted Bytespider because no current official operator documentation was available to verify its purpose and control behavior.
  3. Spot-checked 15 entries (OpenAI, Anthropic, Google, Bing, Perplexity, the three partial/unclear Meta and Amazon rows) against current operator documentation. Corrected Claude-User's robots.txt field from false to true: Anthropic's own bot documentation states plainly that opting out of any of its three bots, including Claude-User, is done through the same robots.txt Disallow mechanism, unlike OpenAI's ChatGPT-User and Perplexity's Perplexity-User, which explicitly say robots.txt rules may not apply to them. No other checked field was wrong.
  4. Re-read the DuckDuckGo, Meta, Semrush, Ahrefs and Majestic crawler pages and added per-bot full user agents, IP-list links, reverse-DNS suffixes, documented notes, and a block-or-allow reading. Corrected three fields: SemrushBot's token is the bare SemrushBot (Semrush documents no version suffix); meta-externalfetcher's robots.txt bypass is for user-requested fetches, not security checks (that reason belongs to FacebookExternalHit); DuckDuckBot's blocking cost now says DuckDuckGo sources most traditional links from Bing.
  5. Re-read the OpenAI, Anthropic, Google and Bing crawler pages and added the same per-bot detail to their ten cards. Corrected Bingbot's reverse-DNS check (Bing documents only search.msn.com, not bing.com) and changed ChatGPT-User's robots.txt field from false to partial, since OpenAI says the rules may not apply, the same wording mapped to partial for Meta's fetcher.
  6. Re-read the Apple, Perplexity, Common Crawl, Amazon, You.com and Diffbot crawler pages and added per-bot detail to their ten cards. Corrected Applebot-Extended, which Apple says does not crawl and only controls training use; Diffbot's token, which Diffbot documents as the bare Diffbot; PerplexityBot's blocking cost, which no longer claims a block removes you from Perplexity answers, since Perplexity-User fetches separately; YouBot's purpose and cost, which You.com describes only in terms of its search results; and Diffbot's robots.txt field, now partial, since Diffbot says robots.txt can be overridden under a customer's agreement with a site.
  7. Re-read the Ai2, ImageSift, Huawei Petal, DataForSEO, Mistral and Liner crawler pages and added per-bot detail. Corrected AI2Bot's robots.txt field to unclear, now labelled not documented (Ai2's notice does not mention robots.txt) and dropped an unsourced dataset name; moved ImagesiftBot to other product crawlers (ImageSift describes a similar-image index, not a public dataset) and fixed its verification field; moved LinerBot to search indexers, as Liner describes it. Timpi documents Collector crawling only in general terms and names no user agent, so the Timpibot card now says so and is kept out of search.
  8. Added three crawlers from their operators' pages: MistralAI-Index and MistralAI-Training, so Mistral's search and training controls each have a card, and OpenAI's OAI-AdsBot, which reviews ChatGPT ad landing pages. OAI-AdsBot's robots.txt field is unclear, because OpenAI does not say whether it reads robots.txt.
  9. Added a publishedRanges field for ranges an operator lists on its bot page rather than in a file (YouBot's /24, DataForSEO's three IPv4 and three IPv6 subnets), and PetalBot's two reverse-DNS domains as a field, so the AI crawler IP verifier can check them.