Meta AI Crawlers Reading the Web While You Block OpenAI

Meta AI Crawlers Reading the Web While You Block OpenAI

TL;DR Summary:

Meta Crawlers Dominate: DataDome says Meta’s crawlers are now the biggest AI readers on the web, with traffic surging far faster than the bots publishers have focused on blocking.

Wrong Bot, Wrong Fight: Many sites are still targeting GPTBot in robots.txt, even though Meta’s crawlers appear to be the ones doing the most reading with little referral traffic in return.

Publishers Miss The Deal: Google triggered backlash because it broke an old traffic-for-access bargain, while Meta never made that promise and has mostly avoided that same public dispute.

Why is Meta’s AI reading your website without paying for it while everyone argues with Google? New data from bot-defense vendor DataDome shows Meta AI crawlers reading the web at a scale that dwarfs the crawlers publishers actually block. This matters right now because the public fight over AI scraping has aimed at the wrong company for months.

What DataDome’s Numbers Show About Meta AI Crawlers Reading the Web

DataDome recorded 17.7 billion AI agent requests across its network in the second quarter of 2026. That is a 45% jump from the first quarter. The growth did not come from Google or OpenAI. Meta-ExternalAgent traffic grew 74% quarter over quarter, and Meta-WebIndexer grew 163%. Together, these two Meta crawlers now carry the majority of the AI agent traffic DataDome tracks. DataDome sells bot protection, so its numbers come from a vendor with a stake in the topic. Still, the pattern is clear enough to take seriously.

GPTBot Gets Blocked, Meta AI Crawlers Reading the Web Get Ignored

Here is the mismatch. GPTBot remains the most-blocked AI crawler in robots.txt files, the text file websites use to tell bots what they can and cannot access. Site owners built their defenses around OpenAI’s bot. Meanwhile, Meta’s crawlers do the heaviest reading and send almost nothing back in referral traffic. Your robots.txt file may be blocking the wrong visitor entirely. This is precisely the blind spot a tool like ClickRank is built to close, giving site owners a clear read on which AI crawlers, Meta’s included, are actually accessing their pages instead of relying on assumptions about which bots deserve blocking.

Why Meta Never Had To Negotiate Like Google Did

Google built a 20-year deal with publishers: Google reads your site, Google sends visitors. When Google’s AI summaries started keeping those visitors instead of sending them to you, publishers felt betrayed. That is why you see blocking, licensing talks, and lawsuits aimed at Google. Meta never made that promise. It has always kept users inside its own apps. So when Meta’s crawlers became the heaviest readers on the open web, no one felt a broken deal, because there was never a deal to break.

The Licensing Deals Meta Signed, and Who They Leave Out

Meta does pay some publishers. In March 2026, it signed a licensing deal with News Corp worth up to $50 million a year, covering both Meta AI answers and model training. Similar deals exist with CNN, Fox News, and USA Today. That covers a handful of the largest media brands. A small blog, a local news site, or a trade publication has no seat at that table and no leverage to get one.

What To Watch Next With Meta AI Crawlers Reading the Web

Reddit can put a licensing deal on its earnings call, as it did on July 30, 2026. Most site owners cannot do that. The real move available to smaller publishers is building direct relationships with readers through owned channels, email lists, and communities that platforms cannot dilute. This takes time and does not run through any negotiation with Google.

Check your server logs, not the headlines. Your logs will tell you which crawlers actually visit your site and how often, and right now that is more likely to be Meta than the AI companies making news. Before you adjust another robots.txt rule based on headlines rather than logs, tools like ClickRank can show you exactly which AI crawlers, Meta’s included, can actually access your pages, so your blocking decisions match what’s really happening rather than what’s making news. The terms the web sets today for Meta AI crawlers reading the web are the terms it will live with when those same crawlers start paying for access instead of taking it for free.


Scroll to Top