Control AI Crawlers With robots.txt, Keep Google
Block AI crawlers by name, one operator token at a time, and keep Googlebot and Bingbot out of those groups. Most AI operators split training, search and user fetches into separate tokens, so you can refuse training and stay in AI search. Three rule sets by goal, the operator split table, and the catches that change what each rule does.
To block AI crawlers with robots.txt without blocking search engines, give each AI crawler its own group, by the exact token its operator documents, and never put Googlebot or Bingbot in those groups. Most AI operators now split their crawling into separate tokens for model training, for their search index, and for fetches a user asks for. That split is what makes a precise policy possible: you can refuse training and stay eligible to be cited in ChatGPT, Claude or Perplexity answers.
Pick your goal, copy the matching rule set below, and test the whole file in our AI crawler robots.txt tester before you deploy it. Every token here comes from its operator's own crawler page, checked on 2026-09-26 and linked from its card in our AI crawler directory.
What robots.txt can and cannot do
Google's robots.txt documentation describes robots.txt as a way to tell crawlers which URLs they can access on your site. It is not access control. Google also notes that a disallowed URL can still be indexed if other pages link to it, even though Google does not crawl it.
For AI crawlers, three more limits matter:
- Some fetchers do not promise to obey it. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, and Meta and Amazon say some of their user-triggered fetches may bypass them. For these, a rule is a stated preference.
- It only covers future crawls. Common Crawl, for example, documents no removal from past crawls; its route for material already collected is a separate opt-out request.
- It is not a legal instrument. If the question is copyright, licensing or privacy, robots.txt is one technical signal next to contracts and access controls, not a substitute for them.
Which operators let you split training from search?
This table shows, operator by operator, which token controls what, so you can see where a clean split is possible and where it is not.
| Operator | Training | Search or AI index | User-requested fetches | The catch |
|---|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User | OpenAI says robots.txt may not apply to ChatGPT-User. |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User | None documented: Anthropic says all three honor robots.txt. |
| Google-Extended (also Gemini grounding) | Googlebot | No token: Google says its user-triggered fetchers ignore robots.txt | Googlebot also feeds AI Overviews and AI Mode. No robots.txt rule keeps you in Search but out of them. | |
| Apple | Applebot-Extended | Applebot | None listed | With no Applebot group, Applebot follows your Googlebot group. |
| Meta | meta-externalagent (also product indexing) | meta-webindexer | meta-externalfetcher | One token covers training and Meta's product indexing. |
| Amazon | Amazonbot (also product improvement) | Amzn-SearchBot | Amzn-User | Amazon also reads the noarchive meta tag as "do not train". |
| Mistral | MistralAI-Training | MistralAI-Index | MistralAI-User | None documented: each token does one job. |
| Perplexity | None: Perplexity says neither agent crawls for foundation models | PerplexityBot | Perplexity-User | Perplexity-User generally ignores robots.txt. |
| DuckDuckGo | None listed | DuckDuckBot for search, DuckAssistBot for AI answers | None listed | A DuckAssistBot block takes 72 hours to apply. |
Read it this way. At OpenAI, Anthropic, Mistral and Amazon, blocking the training token costs nothing in their search products. At Google and Apple, the training token is a control, not a crawler: blocking Google-Extended or Applebot-Extended changes how content is used but does not stop Googlebot or Applebot from crawling. At Meta, refusing training also refuses Meta's product indexing, because one token does both. Meta runs two more crawlers outside this split. Leave facebookexternalhit open on any page you want shared, because it builds the link preview on Facebook, Instagram and Messenger. meta-externalads crawls for Meta's advertising and business products, and Meta documents no referral, preview or citation that depends on it.
Which robots.txt rules match your goal?
Each rule set groups several tokens under one Disallow: /, which the robots.txt standard allows, with a comment naming what each line refuses. Goal 1 ends with a * group that allows everything else. If your file already has a User-agent: * group, keep yours and paste only the named groups. Paste them as their own groups, not under your * line, or they merge with it.
Goal 1: out of AI training, still in AI search
# OpenAI, Anthropic, Mistral model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: MistralAI-Training
# Gemini training and grounding; Googlebot still crawls for Search
User-agent: Google-Extended
# Apple model training; Applebot still crawls for Siri and Spotlight
User-agent: Applebot-Extended
# Meta training (also Meta's product indexing) and Amazon (product use and possible training)
User-agent: meta-externalagent
User-agent: Amazonbot
# Common Crawl's public dataset, which AI trainers download
User-agent: CCBot
# Allen Institute for AI training
User-agent: AI2Bot
Disallow: /
User-agent: *
Allow: /
What it keeps: ChatGPT search, Claude search, Perplexity, Mistral search, Meta AI search, Google and Bing, and user-requested fetches. What it costs: Gemini grounding (Google-Extended covers both training and grounding), Meta's product indexing, Amazon's product-improvement crawl and eligibility for its Content Partners program, which Amazon ties to allowing Amazonbot, and your place in Common Crawl's public datasets and archives. If you want Amazon out of training only, Amazon says it treats the noarchive robots meta tag as "do not use for training", so you can leave Amazonbot crawling and mark the pages instead.
Goal 2: out of AI training and AI answers, still in classic search
Start from Goal 1 and add the AI search indexers and the user-requested fetchers:
# AI search indexes
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: DuckAssistBot
User-agent: meta-webindexer
User-agent: MistralAI-Index
User-agent: Amzn-SearchBot
# Fetches a user asks for (several operators say these may not obey)
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
User-agent: meta-externalfetcher
User-agent: Amzn-User
User-agent: MistralAI-User
Disallow: /
Know its limits before you ship it. Googlebot and Bingbot stay open, so Google's AI Overviews and AI Mode can still use your pages. Google points to snippet controls (nosnippet, data-nosnippet, max-snippet) to limit what its AI features show, and those limit your ordinary snippet too. Apple says the nosnippet tag stops it using a page as context for its AI-generated answers. The user-fetch lines are preferences for the operators who say those fetchers may not obey. And smaller search engines with their own crawlers, such as You.com (YouBot), Liner (LinerBot) and Huawei's Petal (PetalBot, which also feeds Huawei's AI search), stay open; add them if you count them as AI answers.
Goal 3: out of SEO tool databases
User-agent: AhrefsBot
User-agent: SemrushBot
User-agent: MJ12bot
User-agent: DataForSeoBot
Disallow: /
This removes your pages' content and outbound links from these tools; links that other sites point at you stay visible, because the tools find them on those sites. Blocking AhrefsBot also removes you from Yep, the search engine Ahrefs runs. Semrush runs its other tools under separate tokens, such as SiteAuditBot, so this blocks its backlink index only. Each of the four has a directory card with its operator's own terms, including crawl-delay support: AhrefsBot, SemrushBot, MJ12bot and DataForSeoBot.
Step 1: protect search discovery first
If organic search matters, Googlebot and Bingbot must never be collateral damage. Check three things before any AI rule ships:
- important public pages are allowed for Googlebot and Bingbot;
- paths used for rendering, images or structured data are not blocked without a reason;
- no
User-agent: *rule blocks something you meant to block only for AI crawlers: when Googlebot or Bingbot has no group of its own, it follows the*group; - every named group repeats the
*rules it needs, because a crawler with its own group ignores the*group entirely. Bing says so for Bingbot, and the robots.txt standard says the same for every crawler.
That last point catches many sites. If your file has User-agent: * with Disallow: /admin/ and a separate User-agent: Googlebot group, Googlebot does not inherit the /admin/ rule. Repeat it in every named group that needs it.
Step 2: test the whole file before deploying
Paste your current file with the new groups into the robots.txt tester. It runs the file for each crawler in the directory and flags the traps above: a named group that drops your * rules, Applebot or ImagesiftBot following your Googlebot group, control tokens that do not crawl, and fetchers whose operators say they may not obey.
- Test five to ten important public URLs: Googlebot, Bingbot and every crawler you meant to keep should show as allowed.
- Test one URL from each section you meant to block.
- Confirm each blocked token is spelled as its operator documents it.
- Confirm the file is on every host and subdomain you want covered; Anthropic and Semrush both say each subdomain needs its own rule.
- Record the date, the goal, the tokens and the reason in your change log.
- Allow for the operators' delays: OpenAI says about 24 hours for ChatGPT search, DuckDuckGo 72 hours, and Amazon says its bots may use a copy of robots.txt cached for up to 30 days.
Step 3: check that the traffic obeys
After the change has had time to apply, look at your logs. Apart from requests for /robots.txt itself, which every crawler may fetch, a request that still claims to be a blocked crawler after the operator's stated delay is either the real crawler, which would be a problem worth reporting to the operator, or an impostor using its name. Tell them apart with the operator's published IP file or reverse DNS, following our guide to verifying AI crawlers in your server logs. Do not block a crawler's IP ranges as a substitute for robots.txt: Anthropic warns that an IP block can stop its bots from reading your robots.txt at all, so it may not work as an opt-out.
When not to block an AI crawler
- You want your pages cited in that product's answers. Blocking a search or user-fetch token removes that route.
- The block would not be honored anyway. For fetchers whose operators say robots.txt may not apply, a login is the only reliable barrier.
- You cannot verify the current token. Old blog lists carry misspelled or retired names; use the operator's page or the directory card.
- The real question is legal. Robots.txt will not settle copyright, licensing or privacy.
Claim ledger
| Claim | Source | Checked |
|---|---|---|
| robots.txt manages crawler access; a disallowed URL can still be indexed if linked | Google Search Central robots.txt introduction | 2026-09-26 |
A crawler follows the group naming its token and ignores the * group when one exists |
RFC 9309; Google, How Google interprets the robots.txt specification (updated 2026-08-31); Bing, for Bingbot | 2026-09-26 |
| Training, search and user-fetch tokens for OpenAI, Anthropic, Meta, Amazon, Mistral, Perplexity and DuckDuckGo | Each operator's crawler page, via the directory cards | 2026-09-26 |
| Googlebot controls Search including AI features; Google-Extended does not affect Search; snippet controls limit AI features | Google's Googlebot and Google-Extended documentation, via the directory cards | 2026-09-26 |
| Google's user-triggered fetchers ignore robots.txt | Google, Verify requests from Google crawlers and fetchers | 2026-09-26 |
| Applebot follows Googlebot rules when robots.txt names Googlebot but not Applebot; nosnippet keeps content out of Apple's AI answers | Apple's Applebot documentation, via the directory card | 2026-09-26 |
| Amazon reads noarchive as "do not use for training"; its bots may cache robots.txt for up to 30 days | Amazon's Amazonbot documentation, via the directory card | 2026-09-26 |
Review this page when an operator changes its crawler documentation, or by 2026-12-26 with the directory's scheduled recheck.
FAQ
Can robots.txt block AI crawlers?
Yes, for crawlers whose operators say they honor it, when you use the documented token. Most training and search crawlers do. Several user-requested fetchers do not promise to, so for those a rule is a preference.
How do I block AI training but stay in ChatGPT search?
Disallow GPTBot and leave OAI-SearchBot allowed. OpenAI says the two settings are independent. Goal 1 above does the same for the other operators that separate training from search.
How do I block AI bots without blocking Googlebot?
Give each AI crawler its own group by name, and never add Googlebot or Bingbot to those groups or rely on a User-agent: * block. Then test the file for Googlebot and Bingbot on your important URLs.
Does blocking Google-Extended remove me from AI Overviews?
No. Google says Google-Extended does not affect Search, and AI Overviews are part of Search, crawled by Googlebot. It controls Gemini training and grounding.
Should I block ClaudeBot?
Block it if you do not want your content used to train Anthropic's models. It does not affect Claude's search results, which use Claude-SearchBot, or pages a user asks Claude to open, which use Claude-User.
Is robots.txt enough to prevent AI training or legal use of my content?
No. It is a crawler instruction for crawlers that honor it and covers future crawls only. Copyright, licensing and privacy questions need other controls.
Sources
- https://developers.google.com/search/docs/crawling-indexing/robots/intro
- https://developers.google.com/search/docs/crawling-indexing/googlebot
- https://developers.openai.com/api/docs/bots
- https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt
- https://docs.mistral.ai/robots
- https://lastingcontent.com/bots
- https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- https://www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0
- rule sets built from the Lasting Content crawler directory, each token checked against its operator's page on 2026-09-26