Should you block AI2Bot? Ai2's AI training crawler
AI2Bot is operated by Allen Institute for AI. It explores certain domains to find web content, which Ai2 says is used to train open language models.
Who should block AI2Bot, and who should not?
Block it if you do not want your pages in the data Ai2 uses to train open language models. Because the notice does not say the bot reads robots.txt, add the robots.txt rule and also reject the user agent at your server or CDN, which is the method Ai2's own notice points to. Some publishers allow it because the models it trains are open; Ai2 documents no benefit to your site beyond that.
The verified facts
| User-agent token | AI2Bot |
|---|---|
| Full user agent in logs | Mozilla/5.0 (compatible) AI2Bot (+https://www.allenai.org/crawler) |
| Operator | Allen Institute for AI |
| Purpose | Model training crawlers |
| Honors robots.txt | Not documented. The operator does not say whether it honors robots.txt. |
| Verification | No IP list or reverse-DNS method published |
What does Allen Institute for AI’s documentation add?
- Ai2's crawling notice gives the bot's purpose and its user-agent string, and says the string can be used to filter or reject its traffic. It does not mention robots.txt.
- Ai2 publishes no IP list or reverse-DNS method, so the user-agent string is the only documented identifier.
How do you block AI2Bot?
Add this to your robots.txt:
User-agent: AI2Bot
Disallow: /Once you give a crawler its own group, it stops reading your User-agent: * rules, so run the whole file through the AI crawler robots.txt tester before you deploy it.
Because robots.txt compliance is unverified for this bot, a robots.txt line is a request, not an enforcement. Enforcement requires blocking at the CDN or firewall, though without a published IP list that means user-agent matching only.
What does blocking AI2Bot cost you?
Where the block holds, your pages stay out of the content Ai2 collects to train open language models. Ai2's notice does not say whether the bot reads robots.txt.
Which crawlers in the same group should you decide on at the same time?
A robots.txt group for AI2Bot does nothing to crawlers with other tokens. In the same group, the directory also covers MistralAI-Training (Mistral), GPTBot (OpenAI), ClaudeBot (Anthropic) and Applebot-Extended (Apple), each with its own token and its own documented cost of blocking.