Should you block meta-externalagent?

meta-externalagent is operated by Meta. It collects public web content to train Meta's future models and to improve Meta products with indexed content.

Who should block meta-externalagent, and who should not?

Block it if you do not want your content in Meta's model training. The documented cost is small and vague: Meta says the same crawler improves products by indexing content, but names no product that sends readers back. Meta ties citations and links in Meta AI answers to a different crawler, Meta-WebIndexer, so a site can refuse training here and still allow Meta-WebIndexer. Publishers who license their content, or whose archive is the product, have the clearest case for a block.

The verified facts

User-agent tokenmeta-externalagent/1.1
Full user agent in logsmeta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)
OperatorMeta
PurposeModel training and AI grounding controls
Honors robots.txtYes, per the operator's documentation.
VerificationNo official IP ranges published

What does Meta’s documentation add?

  • Meta gives two example uses: training foundation AI models, and improving products by indexing content directly. One token covers both, so you cannot allow the indexing use and refuse the training.
  • Meta's documentation points site owners to robots.txt rather than non-standard formats like NoAI tags. It does not document support for a noai meta tag.

How do you block meta-externalagent?

Add this to your robots.txt:

User-agent: meta-externalagent
Disallow: /

Once you give a crawler its own group, it stops reading your User-agent: * rules, so run the whole file through the AI crawler robots.txt tester before you deploy it.

What does blocking meta-externalagent cost you?

Content excluded from Meta AI training datasets and from product features that use its web index.

What else does Meta document?

Meta also documents meta-externalfetcher, FacebookExternalHit, Meta-WebIndexer and Meta-ExternalAds, each with its own robots.txt token. How Meta’s crawlers fit together.

Which crawlers in the same group should you decide on at the same time?

A robots.txt group for meta-externalagent does nothing to crawlers with other tokens. In the same group, the directory also covers Google-Extended (Google), each with its own token and its own documented cost of blocking.