Meta
Operated by Meta
meta-webindexerAbout this crawler
Meta is Meta's ai search indexing crawler. Indexes the page so it can be surfaced and cited in AI powered search results later. Closest in behaviour to a classic search engine crawler.
A request from Meta usually means the page is being indexed so it can be cited in AI powered search results later, not read by a person at the moment of the request.
User agent
Match on the token, not the whole header. Operators change version numbers and the surrounding text freely, so a substring match on the token is the only stable rule.
Mozilla/5.0 (compatible; meta-webindexer/1.0)
Typical shape. The exact header varies by operator and version.
Blocking it in robots.txt
User-agent: meta-webindexer Disallow: /
robots.txt is a request, not enforcement. Well behaved crawlers honour it; nothing stops one that does not.
Verifying a request really came from Meta
Scrapers routinely borrow a well known crawler's user agent to slip past rate limits. Checking takes three pieces.
1. User agent
What the request claims to be. It is a plain header that anyone can set, so on its own it proves nothing.
2. Request IP
The address the request actually arrived from. Your server sees it directly, and it cannot be set by the caller the way a header can.
3. Range check
This operator publishes neither IP ranges nor a reverse DNS convention. A request carrying this user agent cannot be confirmed as genuine, so treat the count as a claim rather than a fact.
If you need certainty for this one, the practical options are rate limiting by address or blocking it in robots.txt and watching whether the requests stop.
More from Meta
Is Meta allowed on your site?
Zenovay reads your robots.txt and llms.txt and shows, crawler by crawler, which AI products you currently allow and which you block. It also tracks the AI assistants that send you real visitors, and ties those visits to revenue.
Start free