Mistral AI
AI training
MistralAI-Training
Collects web content for datasets that train Mistral's models.
Follows robots.txt- User agent token
- MistralAI-Training
- Operator
- Mistral AI
- Verification
- None published
- Source
- Mistral AI documentation
What an AI training crawler does
A training crawler gathers public pages that may end up in the data used to train AI models. It sends no visitors back, so blocking it is the usual choice for sites that don't want their content used that way.
Block or allow MistralAI-Training in robots.txt
Add one of these to the robots.txt file at the root of your domain. Crawlers read it before they fetch anything else.
Block MistralAI-Training
User-agent: MistralAI-Training
Disallow: /
Allow MistralAI-Training
User-agent: MistralAI-Training
Allow: /
More from Mistral AI
Questions, answered.
MistralAI-Training is an AI training crawler run by Mistral AI. Collects web content for datasets that train Mistral's models.
Yes. Mistral AI says MistralAI-Training follows robots.txt.
Add "User-agent: MistralAI-Training" and "Disallow: /" to your robots.txt.
Mistral AI publishes no IP list for MistralAI-Training, so it can only be recognised by its user agent, which other bots can copy.
No. MistralAI-Training collects pages for model training and sends no visitors back.
Block it if you don't want your pages used to train models. It sends no visitors, so blocking it costs you no traffic.
In its own documentation at https://docs.mistral.ai/robots. Every fact on this page comes from there.
Still have questions? We're happy to help.
See which bots really read your site.
NoirTrack counts MistralAI-Training and every other bot apart from real visitors, and its firewall can block them before they reach your pages.
Start free trialFree 14-day trial, no credit card. The server SDK also sees bots that never run JavaScript.