Should you block ClaudeBot? Anthropic's AI training crawler

ClaudeBot is operated by Anthropic. It collects public web content to train Anthropic's future models. Blocking it keeps your words out of the next training set; it does not remove you from today's AI answers.

Who should block ClaudeBot, and who should not?

Block it if you do not want future material in Anthropic's training data. The block does not remove you from Claude's search or from user-requested fetches, which have their own tokens. If the issue is load rather than training, Crawl-delay is the documented middle ground. Use robots.txt rather than an IP block: Anthropic says an IP block may not hold as an opt-out, because the bot can no longer read your rules.

The verified facts

User-agent tokenClaudeBot
OperatorAnthropic
PurposeModel training crawlers
Honors robots.txtYes, per the operator's documentation.
Published IP rangesclaude.com/crawling/bots.json

What does Anthropic’s documentation add?

  • Anthropic says blocking ClaudeBot signals that the site's future materials should be excluded from its training datasets. The wording covers material going forward.
  • It supports the non-standard Crawl-delay extension, for example Crawl-delay: 1, as a softer alternative to a block.
  • Anthropic asks for the robots.txt rule on every subdomain you want to opt out, and says its bots will not try to bypass CAPTCHAs.

How do you block ClaudeBot?

Add this to your robots.txt:

User-agent: ClaudeBot
Disallow: /

Once you give a crawler its own group, it stops reading your User-agent: * rules, so run the whole file through the AI crawler robots.txt tester before you deploy it.

How do you tell a real ClaudeBot request from a fake one?

A user-agent match proves nothing, since any client can send ClaudeBot. Check the IP against the published range file above, re-fetched on a schedule. A request that fails is not Anthropic. The log verification guide has a script for the IP check and a worked reverse-DNS example, or paste the IP into the AI crawler IP verifier, which runs these checks against Anthropic’s current data.

What does blocking ClaudeBot cost you?

Anthropic says the block signals that the site's future materials should be excluded from its model training datasets. Claude search indexing and user-requested fetches are controlled separately.

What else does Anthropic document?

Anthropic also documents Claude-User and Claude-SearchBot, each with its own robots.txt token. How Anthropic’s crawlers fit together.

Which crawlers in the same group should you decide on at the same time?

A robots.txt group for ClaudeBot does nothing to crawlers with other tokens. In the same group, the directory also covers Applebot-Extended (Apple), AI2Bot (Allen Institute for AI), MistralAI-Training (Mistral) and GPTBot (OpenAI), each with its own token and its own documented cost of blocking.