Free interactive tool
AI Crawler robots.txt Tester
Paste a robots.txt and a path, and the tester reports, for each of the 40 crawlers in our AI crawler directory, whether the file lets it fetch that path, which line decides, and what the crawler’s operator says that a plain verdict hides: when Applebot or ImagesiftBot is following your Googlebot rules, when a crawler’s own group quietly drops the rules you wrote for everyone, when a token such as Google-Extended controls use rather than crawling, and when the operator says the bot may not read robots.txt at all. It runs in your browser and fetches nothing.
How the tester decides
1 Group: the crawler obeys every group that names its token
(case-insensitive; "GPTBot/1.4" counts as GPTBot).
No such group: an operator-documented fallback, if any
(Applebot, ImagesiftBot -> Googlebot). Otherwise the * group.
Otherwise nothing applies and every path is allowed.
2 Rule: of the Allow and Disallow lines in those groups that
match the path, the longest pattern wins. A tie goes to Allow.
"*" matches any characters; a final "$" anchors the end.
3 No matching rule: allowed. /robots.txt: always allowed.
4 Catches: facts from the crawler's own documentation that
change what the verdict means.Sources: steps 1 to 3 follow the Robots Exclusion Protocol standard, RFC 9309 as Google applies it in How Google interprets the robots.txt specification (page last updated 2026-08-31, read 2026-09-26). The tester reproduces every rule-order example on that page. Step 4 comes from each crawler’s card, each read from the operator’s own page and dated there. Most AI operators do not publish how their parser handles edge cases such as wildcards. Where one documents a difference, the tester shows it; otherwise it assumes the standard.
Test your robots.txt
Open https://your-site/robots.txt in a browser and paste the whole file. Each host and subdomain has its own file.
Testing /admin/settings. Query strings count; fragments do not.
Model training crawlers
| Crawler | Verdict | Decided by | What to know |
|---|---|---|---|
| AI2BotAllen Institute for AI | Blocked | the * group: line 3, Disallow: /admin/ |
|
| Applebot-ExtendedApple | Blocked | the * group: line 3, Disallow: /admin/ |
|
| ClaudeBotAnthropic | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| GPTBotOpenAI | Blocked | its own group (GPTBot): line 6, Disallow: / | Nothing beyond the rule. |
| MistralAI-TrainingMistral | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
Model training and AI grounding controls
| Crawler | Verdict | Decided by | What to know |
|---|---|---|---|
| Google-ExtendedGoogle | Blocked | the * group: line 3, Disallow: /admin/ |
|
| meta-externalagentMeta | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
Live retrieval and user-request fetchers
| Crawler | Verdict | Decided by | What to know |
|---|---|---|---|
| Amzn-UserAmazon | Blocked | the * group: line 3, Disallow: /admin/ |
|
| ChatGPT-UserOpenAI | Blocked | the * group: line 3, Disallow: /admin/ |
|
| Claude-UserAnthropic | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| meta-externalfetcherMeta | Blocked | the * group: line 3, Disallow: /admin/ |
|
| MistralAI-UserMistral | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| Perplexity-UserPerplexity | Blocked | the * group: line 3, Disallow: /admin/ |
|
Search and AI-search indexers
| Crawler | Verdict | Decided by | What to know |
|---|---|---|---|
| Amzn-SearchBotAmazon | Blocked | the * group: line 3, Disallow: /admin/ |
|
| ApplebotApple | Allowed | the Googlebot group (operator fallback); no rule matches, so allowed |
|
| BingbotMicrosoft | Allowed | its own group (bingbot): line 12, Allow: / |
|
| Claude-SearchBotAnthropic | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| DiffbotDiffbot | Blocked | the * group: line 3, Disallow: /admin/ |
|
| DuckAssistBotDuckDuckGo | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| DuckDuckBotDuckDuckGo | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| GooglebotGoogle | Allowed | its own group (Googlebot); no rule matches, so allowed |
|
| LinerBotLiner | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| Meta-WebIndexerMeta | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| MistralAI-IndexMistral | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| OAI-SearchBotOpenAI | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| PerplexityBotPerplexity | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| PetalBotHuawei | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| TimpibotTimpi | Blocked | the * group: line 3, Disallow: /admin/ |
|
| YouBotYou.com | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
SEO and marketing tool crawlers
| Crawler | Verdict | Decided by | What to know |
|---|---|---|---|
| AhrefsBotAhrefs | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| DataForSeoBotDataForSEO | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| MJ12botMajestic | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| SemrushBotSemrush | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
Archive and dataset crawlers
| Crawler | Verdict | Decided by | What to know |
|---|---|---|---|
| CCBotCommon Crawl | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
Other product crawlers
| Crawler | Verdict | Decided by | What to know |
|---|---|---|---|
| AmazonbotAmazon | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| FacebookExternalHitMeta | Blocked | the * group: line 3, Disallow: /admin/ |
|
| GoogleOtherGoogle | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| ImagesiftBotImageSift (Hive) | Allowed | the Googlebot group (operator fallback); no rule matches, so allowed |
|
| Meta-ExternalAdsMeta | Blocked | the * group: line 3, Disallow: /admin/ | Nothing beyond the rule. |
| OAI-AdsBotOpenAI | Blocked | the * group: line 3, Disallow: /admin/ |
|
Worked example: one file, four surprises
The tester opens with this hypothetical file, tested against /admin/settings:
User-agent: * Disallow: /admin/ User-agent: GPTBot Disallow: / User-agent: Googlebot Disallow: /search User-agent: bingbot Allow: /
The owner meant to keep every crawler out of /admin/, keep GPTBot out of everything, and keep Google out of internal search results. The file does something else:
- Googlebot may crawl /admin/. It has its own group, so the * group does not apply to it. Only
/searchis closed to it. - Bingbot may crawl /admin/. Same reason: its own group allows everything. Bing’s documentation says so directly.
- Applebot may crawl /admin/ but not /search. There is no Applebot group, and Apple says Applebot then follows the Googlebot rules, not the * rules.
- Google-Extended shows as blocked from /admin/, through the * group. That does not stop a crawl: Googlebot is still allowed. It only tells Google not to use those pages for Gemini.
The fix is to repeat Disallow: /admin/ in the Googlebot and bingbot groups, or to drop those groups if the only thing they add is the /search rule and you are happy to apply it to everyone. Then test /search?q=test and a normal page such as /blog/post to confirm nothing else moved.
What a verdict does not tell you
- Whether your network lets the crawler in. A CDN or firewall bot rule can refuse a crawler that robots.txt allows. OpenAI and Perplexity both ask sites to allow their published IP ranges as well as the token. Anthropic warns that blocking its IPs can stop its bots from reading robots.txt at all.
- When a change takes effect. Google says it generally caches robots.txt for up to 24 hours. OpenAI says about 24 hours for ChatGPT search, DuckDuckGo says 72 hours for DuckAssistBot, and Amazon says its bots may use a copy cached for up to 30 days.
- Whether a page can still appear. Google says a URL blocked by robots.txt can still be indexed if other pages link to it; to keep it out of results, use noindex on a page Google can crawl. To limit what Google’s AI features show from a page, Google points to snippet controls, not robots.txt.
- Which other hosts need the rule. Each host and subdomain serves its own robots.txt, and Anthropic and Semrush both say the rule must be on each subdomain you want covered.
To check that the traffic you see is the crawler it claims to be, use the log verification guide.