> ## Content Index
> Fetch the complete content index at: https://www.dailyferment.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Most answer engine optimization results are not answer engine optimization
- URL: https://www.dailyferment.com/aeo-results-are-mostly-platform-growth/
- Published: 2026-09-19T17:13:10.000Z
- Updated: 2026-09-19T20:35:19.000Z
- Description: The AEO multiples in circulation have no control group. The one study that added one found roughly two thirds of the headline gain was ChatGPT growing, not optimization working.
- Author: Rana Bilal Zafar
- Tags: AI, SEO, #long-argument, #pre-registered

On 17 September I spent an afternoon configuring the answer engine optimization settings on this site. Structured data for LLMs switched on, meta layer written, robots.txt checked so that GPTBot and PerplexityBot would not be turned away at the door. Standard work. Answer engine optimization, which some people call generative engine optimization or GEO, means shaping a page so that ChatGPT, Perplexity and Google's AI summaries quote it rather than skip past it. Then I went looking for evidence that any of it does anything, and what I found has changed how I plan to measure this publication.

The short version: almost every AEO result being quoted right now has no control group. When someone does add one, most of the effect disappears.

## The number everyone is quoting

Search for answer engine optimization statistics and you will hit the same figures within about four clicks. AI referral traffic up 527% year on year. AI search visits up 42.8%. Early adopters capturing 3.4x more AI visibility than late adopters. ChatGPT referrals converting at 14 to 16% against Google organic's 1.76%.

Those numbers are presented as returns on AEO work. Read them again and notice what they actually measure. They measure how much traffic AI platforms sent to websites this year compared with last year. That's a statement about how many people started using ChatGPT, not about whether anyone's optimization worked.

If ChatGPT's user base grows several times over, every site that was already getting a trickle of ChatGPT referrals sees that trickle multiply, including the sites that did nothing at all. Attributing that to your own AEO work is like taking credit for the tide.

## The one study with a control group

In June, Keisuke Watanabe and Kazuki Nakayashiki of Glasp published a [log-based natural experiment](https://arxiv.org/abs/2606.04362?ref=dailyferment.com) that does the obvious thing almost nobody else has done. They applied AEO work to one part of their domain and left the rest of it alone.

The treated corpus was the /youtube/ section of [glasp.co](https://glasp.co/?ref=dailyferment.com), hundreds of thousands of question-and-answer pages. In January 2026 they applied a bundle of four changes: URL canonicalization to consolidate duplicates, demand mining from bot 404 logs to find content gaps, rewriting titles and lead summaries into question-answer format, and a rule protecting already high-performing organic pages from being rewritten. The untreated remainder of the same domain became the control. Same domain authority, same analytics, same bot filtering, same platform tailwind.

Total ChatGPT referrals to the treated pages grew 5.7x. That is the number that would have gone in the case study. The untreated control pages, over the same window, with no work done to them at all, grew 3.5x.

Run a segmented regression on 47 weeks of weekly log-ratio data, 26 before and 21 after, and the level break attributable to the intervention is 1.82x, with a 95% confidence interval of 1.31 to 2.54 and p=0.001\. Filter to engaged sessions only, to strip out bot noise, and it is 2.27x.

What a control group did to the number

Growth in ChatGPT referral traffic, glasp.co, January to June 2026

Reported headline growth on the pages that got the AEO work

5.7x

Same domain, same window, pages left completely alone

3.5x

What was actually left once the control was subtracted

1.82x

Watanabe and Nakayashiki, arXiv:2606.04362\. Whisker shows the 95% confidence interval, 1.31 to 2.54.

The 5.7x is what would have gone in the case study. The 3.5x is what the pages nobody touched did over the same window.

That doesn't mean AEO does nothing. A 1.82x lift is real and worth having. It means roughly two thirds of the headline 5.7x was the platform growing underneath them, and only a third was anything they did.

The authors are more careful about their own finding than the industry quoting them is about theirs. Their placebo-in-time permutation test came back at p=0.16, which they describe as suggestive rather than conclusive, because the pre-intervention period was short and noisy and already trending upward. They list the limitations plainly: single domain, single engine, a bundle of four tactics that cannot be separated from each other, observational rather than randomized. They call for a randomized second-pass rewrite and multi-domain replication before anyone leans on this too hard.

That is one domain, one answer engine, one intervention bundle, by authors who have an interest in AEO working. It's still the most rigorous public evidence on the question, and it's being ignored by every guide currently ranking for the term it describes.

## Where this stopped being an SEO problem

I run engineering delivery. The moment I read the control-group number, I recognised it, because I have been having the same argument all year in a completely different room.

GitKraken surveyed 554 developers and engineering leaders in June. 96.4% of teams had adopted AI coding tools. 84% of developers reported feeling more productive, 43% of them much more productive. And 20% of organisations measured productivity in any specific way. 39% had no way to measure AI's impact at all. Another 33% relied entirely on developers self-reporting that it helped.

The same hole, in a different room

AI and engineering productivity, 554 developers and engineering leaders

Developers who say AI makes them more productive

84%

Organisations that measure it in any specific way

20%

GitKraken, State of AI in Engineering 2026\. A further 39% have no way to measure AI's impact at all.

Same structural error, different industry: a feeling reported by most, measured by almost nobody.

Two different industries, two different vocabularies, the same hole in the middle. Something large and external is moving in each case: AI platform adoption on one side, AI tooling adoption on the other. What moved gets reported as the return on a local decision. And the missing piece is the same in both, though it isn't a better dashboard or a longer report. It's a group that didn't get the treatment.

So here is the claim, as plainly as I can put it. When the whole environment is moving in the direction you want credit for, a measurement without a holdout isn't a weak measurement. It isn't a measurement. It can't tell your work apart from the weather, which means it says nothing about your work.

That lands badly in both places, because a holdout always costs something. Not rewriting half your pages means leaving traffic on the table if the rewrite works. Not giving AI tools to half your engineers means slowing half your engineers down if the tools work, and it's an unpleasant thing to say out loud to the half who get nothing. The cost is real, which is exactly why almost nobody pays it and almost everybody reports numbers anyway.

The one consolation is that the cost is lowest at the start, when you have nothing to lose and no one is watching.

## What I am doing on this site, stated before the results

This publication went live this week with no posts, no index, and no traffic. Its robots.txt currently disallows everything, because it is still private while the content bank gets built. There has never been a cleaner baseline than zero.

So rather than write the AEO case study in nine months, when I'd be free to attribute whatever happened to whatever I felt like, I'm writing the design down now.

Once there is enough content to make it meaningful, roughly half of it will get the full AEO treatment: question-format headings, extractable answer blocks in the opening lines, explicit entity context, FAQ schema where readers genuinely ask the question. The other half will be written to exactly the same editorial standard and will get none of it. Assignment will be by a rule I fix in advance, not by which pieces feel like they deserve it, because letting myself choose is how the better pieces quietly end up in the treated group.

I will report the split, the per-group referral numbers by source, and the difference. If the difference is nothing, that goes up too, with the same prominence. The [rules this publication holds itself to](https://www.dailyferment.com/about/) make that commitment binding rather than optional, which is the main reason for writing them down before having anything to lose by them.

I can already see three ways this will be weaker than the Glasp study. My volume will be far lower, so the noise floor will be higher and small effects will be invisible. A publication where every piece is written to the same standard is a much narrower test than a corpus of hundreds of thousands of pages. And I am one person running a test on my own work, which is the setup most likely to produce a result its author likes. Naming those now is cheaper than discovering them later.

## Marking my confidence

**Established.** The Glasp study's numbers are as reported above. The GitKraken survey's numbers are as reported above. The widely circulated AEO statistics do not control for platform growth, and in the compilations I checked they do not acknowledge the confound at all.

**Inferred.** That the same structural error is operating in AEO reporting and in AI developer productivity reporting. The mechanism is the same and the evidence in both is consistent with it, but nobody has tested the two together and I am reasoning by analogy.

**Guess.** That the real AEO effect, replicated properly across multiple domains and multiple answer engines, lands somewhere between 1.2x and 2x rather than the order-of-magnitude numbers in circulation. One study on one domain is not enough to support that and I would not defend it hard.

What would change my mind: a multi-domain replication with random assignment showing a substantially larger effect, or a good argument that the Glasp control pages were not comparable to the treated ones in some way the authors missed. If either shows up, I will say so here rather than leave this page standing.

## The one thing worth taking from this

Do the AEO work. 1.82x is a genuinely good return for four structural changes, and the downside is close to zero because most of it is just writing more clearly and marking up what you mean.

Just don't confuse it with the tide. Before you quote a multiple at anyone, including yourself, find out what the pages you did nothing to were doing over the same period. If you can't answer that, you don't have a result yet. You have a number.