We block GPTBot but Meta does the reading. Here's the win-one plan.
Publishers are negotiating with Googlebot while Meta's crawlers carry most of the web's AI reading. Google still cites social anyway.
Publishers are negotiating with Googlebot while Meta's crawlers carry most of the web's AI reading. Google still cites social anyway.
We're pointing the machine-access fight at the wrong company. Publishers are deciding whether to block Google's AI or license to it, and Reddit spent part of its July 30 earnings call on exactly that question. But the biggest machine reader of the web belongs to Meta: it reads more than the bots everyone blocks, it sends almost nothing back, and nobody seems to be discussing whether Zuckrawlers should be blocked, too.
This matters because three things changed this month, and they're connected.
DataDome, a bot-defense vendor, reported 17.7 billion AI agent requests across its network in the second quarter of 2026, up 45% from the first. The growth did not come from Google or OpenAI: Meta-ExternalAgent grew 74% quarter over quarter, Meta-WebIndexer grew 163%, and together the two now carry the majority of AI agent traffic. Meanwhile, by the same account, GPTBot remains the most-blocked AI crawler on the web.
That means we organized our defensive walls around the bot that is not doing the reading. It's a moment worth re-examining. Full disclaimer: DataDome sells bot protection, so its telemetry arrives with an interest attached — treat the exact numbers as one network's view. The direction is harder to dismiss.
The industry treats the Google negotiation like a summit, and that's earned. Google pays in traffic — the brand of crawling access. Meta sends close to nothing in referrals (or per DataDome's network-level reports). The access-control logic in the site operator's mind should differ depending on whether a crawler gives something back.
Hmm — that's the framing that matters. "Every move at that table gets covered like a summit" — and meanwhile, Meta reads freely, unnoticed.
Independently, Google is rolling out the August 2026 spam update, the third of the year after March and June rollouts. What's notable is what's absent: Google didn't announce new spam policies or a companion blog post. Its spam policies page hasn't changed since December. This is pure algorithm enforcement, with no new "rules" to study, it's a remix of existing walls. For marketing teams, the move isn't to read a policy; it's to check whether you actually follow the existing ones — especially autogenerated content, "scaled content abuse," and content that doesn't meet the page's purpose.
While the team blocks bots, the AI answers are busy citing social content at scale. Analysis of over 300 million U.S. monthly searches shows Facebook was cited in 19.5 million AI Overviews, Instagram in 877,000, TikTok in 78,000 — and about one in 15 U.S. searches now puts Facebook or Instagram content inside a Google AI answer.
The critical insight: AI doesn't reward for large following. A 5,000-follower account with real expertise or hard numbers "can beat a 500,000-follower brand account that posts filler". Reach is not influence. What AI rewards is the content that answers the exact question in front of it.
1. Audit your robots.txt against reality. Check whether Meta-ExternalAgent and Meta-WebIndexer are treated the way GPTBot is. Decide honestly, not based on panic, whether these crawlers return enough referral signal to keep open — or whether you're funneling the majority of AI reading for zero return.
2. Run a content-admin audit of soft definitions. If Google's enforcement tightens around spam without new policy, your scaled-automation costs will rise. Trim anything that doesn't measurably serve intent.
3. Re-label social as a) a citation source, not a platform. Cross-link from Facebook and Instagram to actual answers on your domain. The IIO era is an answer-market, and a small, precise profile beats a big, blurry one. That's the new mover.
4. Stop negotiating with bots and start writing for the answer. The same over the circumstances — the cited excerpt wins, not the biggest account.
We're defending against the wrong machine. The bot the web organized around is not the bot doing the reading. The move now is simpler: answer questions well, in public, on the web, in shareable — and stop the argue about GPTBot. Fake wall — it's a dent in the ground.
Verified live · all four AEO engines + the four major web indexes · last reviewed Jun 27, 2026