JOURNAL  /  AEO

Main Street moved to AI answers, the engines closed the hood, and nobody built a scoreboard.

OpenAI onboards small business, hardens its models, and voice evaluation lags the flood. The GEO read — and the move to make this week.

3 MIN READ 617 WORDS
Main Street moved to AI answers, the engines closed the hood, and nobody built a scoreboard.
FIG. 01 — GEO Operating Conditions: Distribution, Opacity, Measurement

Three signals, one story

Generative engine optimization rarely gets breaking news. This week it got three headlines that look unrelated and aren't. OpenAI is partnering with America's SBDC to expand hands-on AI training and local support for small businesses, alongside a new report on how small teams actually use AI. Its security team disrupted a coordinated campaign to extract protected model reasoning and is hardening defenses against adversarial distillation. And Hugging Face shipped an open leaderboard to standardize evaluation for multilingual text-to-speech and voice cloning, because the field now has more than 8,000 TTS models on the Hub as of September 30, 2026 and no agreed way to grade them.

Distribution is widening. The models are closing. Measurement is lagging. That's the operating environment for GEO right now.

Main Street is an AI-first buying surface now

When AI training shows up at the local SBDC office, the buyer journey changes shape. The behavior that B2B marketers have spent two years adapting to — asking a model first, trusting the answer, rarely clicking — is being taught, deliberately, to the smallest businesses. The practical consequence is uncomfortable for anyone still treating AI answers as an enterprise phenomenon: local and SMB commercial queries increasingly get resolved inside a generative engine, and the cited source wins the consideration set without a click. If small businesses are in your pipeline — or you are one — your visibility in AI answers is a storefront question, not an experiment.

The hood is closing. Stop trying to reverse-engineer it.

The distillation story reads like security news. Read it as a transparency forecast instead: a major lab treats its model's reasoning as protected property and will spend real resources defending it. That tells you the "why" behind any engine's outputs is going to stay opaque — no leaked ranking formula, no reverse-engineered prompt secret. It should also raise your guardrails against vendors selling proprietary insight into "how the model thinks." GEO's method is empirical and always was: run a fixed set of prompts, log what comes back, change the inputs, measure the delta. You optimize the answer surface you can observe, not the mechanism you can't.

There is no MOS for citations

The TTS situation is the cleanest analogy marketers currently have for their own predicament. Voice generation has thousands of models, but the gold standard remains human preference scores like MOS or MUSHRA, and arena leaderboards — the scalable proxy — can't keep up. As of late September, only 16 of the 92 models on Artificial Analysis are open-weights, partly because adding an API model takes an API key while an open model must be hosted and served by the arena operator. And even where arenas exist, they struggle to keep voters consistent.

Swap the nouns. Generative engines run on a shifting mix of thousands of models. Deployment changes monthly. And there is no standardized, human-anchored metric for "which sources get cited." Most commercial "AI visibility scores" are arenas with a tiny, unmonitored electorate. The honest position: nobody's vendor score is the ground truth, so build your own. Fixed prompt set. Weekly runs across ChatGPT, Perplexity, and AI Overviews. Citation counts by engine. Share of voice in your category. That spreadsheet is your leaderboard, and unlike the arenas, you control the electorate.

Voice is the second citation surface

Eight thousand TTS models and cheap voice cloning mean spoken answers are about to scale the way text answers did — across languages, not just English. A voice assistant that reads an answer aloud picks fewer sources than a page that can list ten links. Being cited in text is table stakes; being the one source a voice answer reads from is a much narrower cut. Start treating your most-cited pages as voice

§ DISCOVERABLE ON

This domain is surfaced in every channel we ship for clients.

Verified live · all four AEO engines + the four major web indexes · last reviewed Jun 27, 2026

Chat on WhatsApp