Should you block MistralAI-Training?
MistralAI-Training is operated by Mistral. It crawls web content to build datasets for training Mistral's generative AI models. Blocking it keeps your pages out of those datasets; it does not remove you from Mistral search or Vibe answers.
Who should block MistralAI-Training, and who should not?
Block it if you do not want Mistral training on your writing. Mistral separates training from search and live fetches, so the block costs nothing in Vibe answers or Mistral search. That makes it the cheapest Mistral rule, like GPTBot for OpenAI. The limit: with no published addresses, a request that claims to be MistralAI-Training cannot be confirmed as Mistral's, so the robots.txt rule is the lever, not a firewall list.
The verified facts
| User-agent token | MistralAI-Training/1.0 |
|---|---|
| Full user agent in logs | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots) |
| Operator | Mistral |
| Purpose | Model training crawlers |
| Honors robots.txt | Yes, per the operator's documentation. |
| Verification | No IP list published: Mistral publishes address files for MistralAI-User and MistralAI-Index, not for this crawler |
What does Mistral’s documentation add?
- Mistral says webmasters can disallow MistralAI-Training in robots.txt, and that it is not used for search indexing or to answer live user queries in Vibe.
- Its section of Mistral's page gives a user-agent string but, unlike the two other Mistral agents, no IP file.
How do you block MistralAI-Training?
Add this to your robots.txt:
User-agent: MistralAI-Training
Disallow: /Once you give a crawler its own group, it stops reading your User-agent: * rules, so run the whole file through the AI crawler robots.txt tester before you deploy it.
What does blocking MistralAI-Training cost you?
Your content is kept out of the datasets Mistral builds to train its models. Mistral says this crawler serves no search index and no live answers.
What else does Mistral document?
Mistral also documents MistralAI-User and MistralAI-Index, each with its own robots.txt token. How Mistral’s crawlers fit together.
Which crawlers in the same group should you decide on at the same time?
A robots.txt group for MistralAI-Training does nothing to crawlers with other tokens. In the same group, the directory also covers GPTBot (OpenAI), ClaudeBot (Anthropic), Applebot-Extended (Apple) and AI2Bot (Allen Institute for AI), each with its own token and its own documented cost of blocking.