Should you block Meta-WebIndexer? Meta's AI search crawler

Meta-WebIndexer is operated by Meta. It crawls and indexes web content to improve Meta AI search results, citations, and links.

Who should block Meta-WebIndexer, and who should not?

Allow it if you want Meta AI answers to link to you; it is the one Meta crawler whose documented job includes citation. A site that wants to stay out of Meta's model training and keep those citations can block meta-externalagent and leave Meta-WebIndexer open. Block it if you do not want Meta AI to index your pages for answers, knowing the block also gives up the links. It does not stop meta-externalfetcher, which fetches pages at a user's request and may bypass robots.txt.

The verified facts

User-agent tokenmeta-webindexer/1.1
Full user agent in logsmeta-webindexer/1.1 (+/documentation/sharing/webmasters/web-crawlers)
OperatorMeta
PurposeSearch and AI-search indexers
Honors robots.txtYes, per the operator's documentation.
VerificationNo official IP ranges published

What does Meta’s documentation add?

  • Meta says allowing it helps Meta cite and link to your content in Meta AI's responses. It is the Meta crawler tied to citations.
  • Meta describes its purpose as improving Meta AI search result quality. Unlike meta-externalagent, its description does not mention model training.

How do you block Meta-WebIndexer?

Add this to your robots.txt:

User-agent: meta-webindexer
Disallow: /

Once you give a crawler its own group, it stops reading your User-agent: * rules, so run the whole file through the AI crawler robots.txt tester before you deploy it.

What does blocking Meta-WebIndexer cost you?

Content may not be cited or linked in Meta AI search responses.

What else does Meta document?

Meta also documents meta-externalagent, meta-externalfetcher, FacebookExternalHit and Meta-ExternalAds, each with its own robots.txt token. How Meta’s crawlers fit together.

Which crawlers in the same group should you decide on at the same time?

A robots.txt group for meta-webindexer does nothing to crawlers with other tokens. In the same group, the directory also covers Amzn-SearchBot (Amazon), YouBot (You.com), Diffbot (Diffbot), PetalBot (Huawei), MistralAI-Index (Mistral) and LinerBot (Liner), each with its own token and its own documented cost of blocking.