Should you block meta-externalagent?
meta-externalagent is operated by Meta. It collects public web content to train Meta's future models and to improve Meta products with indexed content.
Who should block meta-externalagent, and who should not?
Block it if you do not want your content in Meta's model training. The documented cost is small and vague: Meta says the same crawler improves products by indexing content, but names no product that sends readers back. Meta ties citations and links in Meta AI answers to a different crawler, Meta-WebIndexer, so a site can refuse training here and still allow Meta-WebIndexer. Publishers who license their content, or whose archive is the product, have the clearest case for a block.
The verified facts
| User-agent token | meta-externalagent/1.1 |
|---|---|
| Full user agent in logs | meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers) |
| Operator | Meta |
| Purpose | Model training and AI grounding controls |
| Honors robots.txt | Yes, per the operator's documentation. |
| Verification | No official IP ranges published |
What does Meta’s documentation add?
- Meta gives two example uses: training foundation AI models, and improving products by indexing content directly. One token covers both, so you cannot allow the indexing use and refuse the training.
- Meta's documentation points site owners to robots.txt rather than non-standard formats like NoAI tags. It does not document support for a noai meta tag.
How do you block meta-externalagent?
Add this to your robots.txt:
User-agent: meta-externalagent
Disallow: /Once you give a crawler its own group, it stops reading your User-agent: * rules, so run the whole file through the AI crawler robots.txt tester before you deploy it.
What does blocking meta-externalagent cost you?
Content excluded from Meta AI training datasets and from product features that use its web index.
What else does Meta document?
Meta also documents meta-externalfetcher, FacebookExternalHit, Meta-WebIndexer and Meta-ExternalAds, each with its own robots.txt token. How Meta’s crawlers fit together.
Which crawlers in the same group should you decide on at the same time?
A robots.txt group for meta-externalagent does nothing to crawlers with other tokens. In the same group, the directory also covers Google-Extended (Google), each with its own token and its own documented cost of blocking.