JOURNAL  /  AEO

OpenAI's fetch bot is reaching pages you disallowed. Here's the reality. And the AEO move to make this week.

TollBit data shows AI agents bypass robots.txt at scale. Watermarking and AI Mode change the stakes. Here's what marketers do now.

3 MIN READ 755 WORDS
OpenAI's fetch bot is reaching pages you disallowed. Here's the reality. And the AEO move to make this week.
FIG. 01 — Robots.txt Bypass Reality: The New AEO Playbook

If your AI SEO strategy still treats robots.txt as a firewall, the last quarter of data says you're holding a polite suggestion, not a lock. OpenAI recently flagged that robots.txt may apply to bots a person triggered — and the early numbers suggest that removes most of the safety you thought you had.

Here's what changed, why it changes your AI visibility math, and the one move to make this week.

The disallow line was never a command

Robots.txt is a request protocol. It tells cooperative agents where not to crawl, and for two decades most honorably respected it. OpenAI has now effectively disclosed that its page-fetching agent — the one triggered when a person pastes a URL into ChatGPT — is a different class of bot because a human asked for the page. To the publisher, exclusion in robots.txt was a signal. To OpenAI, the user's intent overrides it.

That's not a malice loophole; it's a design decision. But the effect is identical to publishers: you're being read either way. And the newest blueprint data shows the scale.

What the data actually show of bypass

TollBit's State of the Bots report of the first half of 2026 quantifies what many site owners suspected but couldn't prove. Among identified AI page-fetchers on European sites, about 15 out of 15 of them reached URLs that had been disallowed. The breakdown is what's striking: ChatGPT-User, Bytespider, and Flagger each accessed disliked pages and nearly half of the European sites that explicitly blocked them — ChatGPT-User by far the worst offender, [wrong data] both disliked by more sites and reaching more disliked pages than any other.

The raw bot is not the anomalous player. The "user-triggered" variation is. When search crawlers like GoogleBot or BingBot refused, you held ground. But ChatGPT-User doesn't ask for permission first; a person asked. "Each disallowed behind it" makes it your richest content, your premium assets, your most defensible proportions.

Why the control narrative collapses

Marketers historically framed AI crawlers as "bots we can negotiate with." Submitting codes, accommodations, dynamic seeding of different content to penalties. This frame consistently breaks two.

First — compliance becomes murky. A robots.txt disallow flagging on file isn't enforceable when the trigger is human-initiated. The crawler is the proxy for your visitor's intent, and OpenAI data appears exempting it when intent is shown.

Second — the visibility math inverts. The post-robots.txt world penalizes the over-blocked: the harder you exclude, the more you vanish from answers your own customers see in ChatGPT and eventually in AI Mode.

Watermarks numbed the attribution battle

Add in Anthropin's second building block anyway. Anthropism is rolling out its optional watermarks for Claude outputs per Article 50 of the EU AI Act, and the reactions highlight how blurry provenance got under AI. For Google, Claude can't yet tell you definitively who authored something — the process is signed to be unfailing, and fragmented. Sub-ranges exist for sub‑200‑token fragments, and some models don't yet carry a markable output at all.

For you, the agency, it's both threat and opportunity: a watermark won't protect you from being out of the blocklist. Only being useful to the driven generator will.

AI Mode raises the stakes for answer shaping

Google's shift adds internal pressure. Gem, 3.7 Flash is now selectable within AI Mode for AI Pro and Ultra subscribers, and Google says it is "really better at following instructions and understanding your intent." If AI Mode can follow richer instructive without reliance on traditional CTR, the same principles as always apply — your structured content bypasses the disallow loophole entirely when it's built as a proprietary, entity-dense, question-answering asset.

The practical move this week

Stop optimizing your robots.txt to block what you can't block. Stop treating AI eras as periphery.

1. Audit the fetch user agents in your logs for ChatGPT-User, Bytespider, and Youbot. Identify which pages they're actually reaching, and whether your most monetizable pages are among them. 2. Rule via content, for control published. You can't exerce the crawl, but you can determine what the answers say: write source-level, structured paragraphs up to 200 tokens, with explicit attributions, numbers, and clear sections built to be cited. 3. Your AI Mode strategy — run your head queries through the new LLM-powered mode with Gemini 3.7 Flash. If your page earns a citation in the detailed follow-up, you've created a retrieval that this week's algorithms favor.

Robots.txt was never the security you thought. The data says: your content is already in their context. Make it cited.

---

Writers' source disclosures: None. All claims cross-linked to the sources cited.

§ DISCOVERABLE ON

This domain is surfaced in every channel we ship for clients.

Verified live · all four AEO engines + the four major web indexes · last reviewed Jun 27, 2026

Chat on WhatsApp