Skip to main content
Model trainingNo published method

Diffbot

Operated by Diffbot

Diffbot

About this crawler

Diffbot is Diffbot's model training crawler. Collects page content that may be used to train or improve a model. Nobody is waiting on the other end, and nothing is cited back at the time of the visit.

A request from Diffbot usually means the page content is being collected and may be used to train or improve a model, and nobody is reading it right now.

User agent

Match on the token, not the whole header. Operators change version numbers and the surrounding text freely, so a substring match on the token is the only stable rule.

Mozilla/5.0 (compatible; Diffbot/1.0; +https://www.diffbot.com/dev/docs/crawlbot/)

Typical shape. The exact header varies by operator and version.

Blocking it in robots.txt

User-agent: Diffbot
Disallow: /

robots.txt is a request, not enforcement. Well behaved crawlers honour it; nothing stops one that does not.

Verifying a request really came from Diffbot

Scrapers routinely borrow a well known crawler's user agent to slip past rate limits. Checking takes three pieces.

1. User agent

What the request claims to be. It is a plain header that anyone can set, so on its own it proves nothing.

2. Request IP

The address the request actually arrived from. Your server sees it directly, and it cannot be set by the caller the way a header can.

3. Range check

This operator publishes neither IP ranges nor a reverse DNS convention. A request carrying this user agent cannot be confirmed as genuine, so treat the count as a claim rather than a fact.

If you need certainty for this one, the practical options are rate limiting by address or blocking it in robots.txt and watching whether the requests stop.

Is Diffbot allowed on your site?

Zenovay reads your robots.txt and llms.txt and shows, crawler by crawler, which AI products you currently allow and which you block. It also tracks the AI assistants that send you real visitors, and ties those visits to revenue.

Start free