Free interactive tool

AI Crawler robots.txt Tester

Paste a robots.txt and a path, and the tester reports, for each of the 40 crawlers in our AI crawler directory, whether the file lets it fetch that path, which line decides, and what the crawler’s operator says that a plain verdict hides: when Applebot or ImagesiftBot is following your Googlebot rules, when a crawler’s own group quietly drops the rules you wrote for everyone, when a token such as Google-Extended controls use rather than crawling, and when the operator says the bot may not read robots.txt at all. It runs in your browser and fetches nothing.

How the tester decides

1  Group: the crawler obeys every group that names its token
     (case-insensitive; "GPTBot/1.4" counts as GPTBot).
     No such group: an operator-documented fallback, if any
     (Applebot, ImagesiftBot -> Googlebot). Otherwise the * group.
     Otherwise nothing applies and every path is allowed.
2  Rule: of the Allow and Disallow lines in those groups that
     match the path, the longest pattern wins. A tie goes to Allow.
     "*" matches any characters; a final "$" anchors the end.
3  No matching rule: allowed. /robots.txt: always allowed.
4  Catches: facts from the crawler's own documentation that
     change what the verdict means.

Sources: steps 1 to 3 follow the Robots Exclusion Protocol standard, RFC 9309 as Google applies it in How Google interprets the robots.txt specification (page last updated 2026-08-31, read 2026-09-26). The tester reproduces every rule-order example on that page. Step 4 comes from each crawler’s card, each read from the operator’s own page and dated there. Most AI operators do not publish how their parser handles edge cases such as wildcards. Where one documents a difference, the tester shows it; otherwise it assumes the standard.

Test your robots.txt

Start from:

Open https://your-site/robots.txt in a browser and paste the whole file. Each host and subdomain has its own file.

Testing /admin/settings. Query strings count; fragments do not.

Show
Result for /admin/settings36 of 40 crawlers are asked not to fetch this path. Model training crawlers: 5 of 5 blocked. Model training and AI grounding controls: 2 of 2 blocked. Live retrieval and user-request fetchers: 6 of 6 blocked. Search and AI-search indexers: 13 of 16 blocked. SEO and marketing tool crawlers: 4 of 4 blocked. Archive and dataset crawlers: 1 of 1 blocked. Other product crawlers: 5 of 6 blocked.

Model training crawlers

CrawlerVerdictDecided byWhat to know
AI2BotAllen Institute for AIBlockedthe * group: line 3, Disallow: /admin/
  • Allen Institute for AI does not say whether AI2Bot reads robots.txt.
Applebot-ExtendedAppleBlockedthe * group: line 3, Disallow: /admin/
  • Applebot-Extended does not crawl. Whether these URLs are crawled depends on the Applebot verdict; this one only decides whether Apple may use them to train its foundation models.
ClaudeBotAnthropicBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
GPTBotOpenAIBlockedits own group (GPTBot): line 6, Disallow: /Nothing beyond the rule.
MistralAI-TrainingMistralBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.

Model training and AI grounding controls

CrawlerVerdictDecided byWhat to know
Google-ExtendedGoogleBlockedthe * group: line 3, Disallow: /admin/
  • Google-Extended fetches nothing itself. Whether these URLs are crawled depends on the Googlebot verdict; this one only decides whether Google may use them to train or ground Gemini.
meta-externalagentMetaBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.

Live retrieval and user-request fetchers

CrawlerVerdictDecided byWhat to know
Amzn-UserAmazonBlockedthe * group: line 3, Disallow: /admin/
  • Amazon says some Amzn-User requests may not follow robots.txt. Treat a block as a stated preference.
ChatGPT-UserOpenAIBlockedthe * group: line 3, Disallow: /admin/
  • OpenAI says some ChatGPT-User requests may not follow robots.txt. Treat a block as a stated preference.
Claude-UserAnthropicBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
meta-externalfetcherMetaBlockedthe * group: line 3, Disallow: /admin/
  • Meta says some meta-externalfetcher requests may not follow robots.txt. Treat a block as a stated preference.
MistralAI-UserMistralBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
Perplexity-UserPerplexityBlockedthe * group: line 3, Disallow: /admin/
  • Perplexity says Perplexity-User generally ignores robots.txt, because a person asked for the page. The verdict is what your file asks for, not what will happen.

Search and AI-search indexers

CrawlerVerdictDecided byWhat to know
Amzn-SearchBotAmazonBlockedthe * group: line 3, Disallow: /admin/
  • Amazon says that when robots.txt does not mention Amzn-SearchBot but allows other search bots, it crawls under the rules given to those bots. Name it to be sure.
ApplebotAppleAllowedthe Googlebot group (operator fallback); no rule matches, so allowed
  • Your file has no Applebot group but does name Googlebot, and Apple says Applebot then follows the Googlebot rules.
  • The * group's "Disallow: /admin/" (line 3) would block this path, but Applebot reads the Googlebot group it follows instead of the * group. Repeat the rule there if you meant it for every crawler.
BingbotMicrosoftAllowedits own group (bingbot): line 12, Allow: /
  • The * group's "Disallow: /admin/" (line 3) would block this path, but Bingbot reads its own group instead of the * group. Repeat the rule there if you meant it for every crawler.
Claude-SearchBotAnthropicBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
DiffbotDiffbotBlockedthe * group: line 3, Disallow: /admin/
  • Diffbot says some Diffbot requests may not follow robots.txt. Treat a block as a stated preference.
DuckAssistBotDuckDuckGoBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
DuckDuckBotDuckDuckGoBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
GooglebotGoogleAllowedits own group (Googlebot); no rule matches, so allowed
  • The * group's "Disallow: /admin/" (line 3) would block this path, but Googlebot reads its own group instead of the * group. Repeat the rule there if you meant it for every crawler.
LinerBotLinerBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
Meta-WebIndexerMetaBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
MistralAI-IndexMistralBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
OAI-SearchBotOpenAIBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
PerplexityBotPerplexityBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
PetalBotHuaweiBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
TimpibotTimpiBlockedthe * group: line 3, Disallow: /admin/
  • Timpi does not say whether Timpibot reads robots.txt.
YouBotYou.comBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.

SEO and marketing tool crawlers

CrawlerVerdictDecided byWhat to know
AhrefsBotAhrefsBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
DataForSeoBotDataForSEOBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
MJ12botMajesticBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
SemrushBotSemrushBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.

Archive and dataset crawlers

CrawlerVerdictDecided byWhat to know
CCBotCommon CrawlBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.

Other product crawlers

CrawlerVerdictDecided byWhat to know
AmazonbotAmazonBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
FacebookExternalHitMetaBlockedthe * group: line 3, Disallow: /admin/
  • Meta says some FacebookExternalHit requests may not follow robots.txt. Treat a block as a stated preference.
GoogleOtherGoogleBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
ImagesiftBotImageSift (Hive)Allowedthe Googlebot group (operator fallback); no rule matches, so allowed
  • Your file has no ImagesiftBot group but does name Googlebot, and ImageSift says ImagesiftBot then follows the Googlebot rules.
  • The * group's "Disallow: /admin/" (line 3) would block this path, but ImagesiftBot reads the Googlebot group it follows instead of the * group. Repeat the rule there if you meant it for every crawler.
Meta-ExternalAdsMetaBlockedthe * group: line 3, Disallow: /admin/Nothing beyond the rule.
OAI-AdsBotOpenAIBlockedthe * group: line 3, Disallow: /admin/
  • OpenAI does not say whether OAI-AdsBot reads robots.txt.

Worked example: one file, four surprises

The tester opens with this hypothetical file, tested against /admin/settings:

User-agent: *
Disallow: /admin/

User-agent: GPTBot
Disallow: /

User-agent: Googlebot
Disallow: /search

User-agent: bingbot
Allow: /

The owner meant to keep every crawler out of /admin/, keep GPTBot out of everything, and keep Google out of internal search results. The file does something else:

  • Googlebot may crawl /admin/. It has its own group, so the * group does not apply to it. Only /search is closed to it.
  • Bingbot may crawl /admin/. Same reason: its own group allows everything. Bing’s documentation says so directly.
  • Applebot may crawl /admin/ but not /search. There is no Applebot group, and Apple says Applebot then follows the Googlebot rules, not the * rules.
  • Google-Extended shows as blocked from /admin/, through the * group. That does not stop a crawl: Googlebot is still allowed. It only tells Google not to use those pages for Gemini.

The fix is to repeat Disallow: /admin/ in the Googlebot and bingbot groups, or to drop those groups if the only thing they add is the /search rule and you are happy to apply it to everyone. Then test /search?q=test and a normal page such as /blog/post to confirm nothing else moved.

What a verdict does not tell you

  • Whether your network lets the crawler in. A CDN or firewall bot rule can refuse a crawler that robots.txt allows. OpenAI and Perplexity both ask sites to allow their published IP ranges as well as the token. Anthropic warns that blocking its IPs can stop its bots from reading robots.txt at all.
  • When a change takes effect. Google says it generally caches robots.txt for up to 24 hours. OpenAI says about 24 hours for ChatGPT search, DuckDuckGo says 72 hours for DuckAssistBot, and Amazon says its bots may use a copy cached for up to 30 days.
  • Whether a page can still appear. Google says a URL blocked by robots.txt can still be indexed if other pages link to it; to keep it out of results, use noindex on a page Google can crawl. To limit what Google’s AI features show from a page, Google points to snippet controls, not robots.txt.
  • Which other hosts need the rule. Each host and subdomain serves its own robots.txt, and Anthropic and Semrush both say the rule must be on each subdomain you want covered.

To check that the traffic you see is the crawler it claims to be, use the log verification guide.