Should you block Diffbot? Diffbot's knowledge-graph crawler
Diffbot is operated by Diffbot. It proactively crawls the web for Diffbot's general search engine and Knowledge Graph. Diffbot says it is not used for AI training.
Who should block Diffbot, and who should not?
Blocking the Diffbot token is low-cost for most sites: it removes you from a knowledge graph and search index that people reach through Diffbot's APIs rather than by browsing. Two limits to know. The block does not cover Diffbot-User, which has its own token. And crawls that Diffbot's customers run through its software may use their own user agent, so a Diffbot rule will not catch them. Block both tokens if you want Diffbot's own traffic out.
The verified facts
| User-agent token | Diffbot |
|---|---|
| Operator | Diffbot |
| Purpose | Search and AI-search indexers |
| Honors robots.txt | Partly. The operator says that some requests may bypass robots.txt. |
| Verification | No official IP ranges published |
What does Diffbot’s documentation add?
- Diffbot says its crawler builds a general search engine and Knowledge Graph and is not used for AI training.
- It adheres to robots.txt, including Crawl-delay, by default. Diffbot says robots.txt can be overridden in specific cases, typically when its customer has an agreement with the site being crawled.
- A second agent, Diffbot-User, handles requests a person makes through Diffbot software. It has its own robots.txt token.
- Diffbot's customers can also run their own crawls through its Crawlbot and Extract products, where Diffbot recommends they set their own user agent and robots.txt adherence is a setting that is on by default and can be switched off. Those requests may not say Diffbot at all.
How do you block Diffbot?
Add this to your robots.txt:
User-agent: Diffbot
Disallow: /Once you give a crawler its own group, it stops reading your User-agent: * rules, so run the whole file through the AI crawler robots.txt tester before you deploy it.
Because this bot may bypass some robots.txt directives, the rule does not enforce every request. Enforcement requires blocking at the CDN or firewall, though without a published IP list that means user-agent matching only.
What does blocking Diffbot cost you?
Your pages drop out of the web index behind the Diffbot Knowledge Graph and its search services.
Which crawlers in the same group should you decide on at the same time?
A robots.txt group for Diffbot does nothing to crawlers with other tokens. In the same group, the directory also covers PetalBot (Huawei), MistralAI-Index (Mistral), LinerBot (Liner), OAI-SearchBot (OpenAI), Claude-SearchBot (Anthropic) and Googlebot (Google), each with its own token and its own documented cost of blocking.