What Is GEO? What the Evidence Actually Says About AI Citations

July 28, 2026 — 12 min read — Discoverability & Content

Generative engine optimization (GEO) is the practice of structuring content so that AI answer engines — ChatGPT, Perplexity, Google's AI Overviews and AI Mode — cite it inside the answers they generate. The goal shifts from occupying a position in a list of links to being named and quoted inside a synthesized answer. That much is uncontroversial. What comes next mostly isn't.

If you have read anything about GEO in the last two years, you have met the number: optimizing for AI increases visibility by up to 40%. It comes from a real, peer-reviewed paper. It also measured something considerably narrower than the way it gets quoted — and a July 2026 survey reviewing 45 studies lists the general version of that claim in its table of findings under a single word: rejected.

This post is what the evidence actually supports, what it doesn't, and what a Seed–Series B team should do on Monday.

Where the "40% more visibility" number came from

The foundational work is GEO: Generative Engine Optimization by Aggarwal and colleagues, published at ACM KDD 2024. It named the problem and built the first benchmark — 10,000 queries, with nine content strategies tested against a control.

For each query, the top five Google results were placed into the model's context, one of them was rewritten under a given strategy, and the answer was measured for how much attributed text each source received. Adding quotations moved the position-adjusted word count from 19.3 to 27.2 — a relative gain of about 41%. Keyword stuffing made things worse.

So the finding is real, and it is this: a document already sitting in the model's context can be rewritten to claim a larger share of the answer. It is not a finding about whether your page gets retrieved in the first place. No clicks, referrals or conversions were observed anywhere in the study.

Figure 1. The 2024 GEO benchmark held retrieval, ranking and click-through fixed and measured only the generation-and-citation stage.
Where the "+40% AI visibility" figure actually applies The AI citation pipeline has six stages: activation, crawling and indexing, retrieval, reranking and context allocation, generation and citation, and click or conversion. The 2024 GEO benchmark held the first four stages fixed by placing five documents directly into the model's context. It measured only the generation and citation stage, where adding quotations raised a source's share of attributed text by about 41 percent. Clicks, referrals and conversions were never observed. Where the “+40% AI visibility” figure applies The citation pipeline, and the one stage the 2024 benchmark actually measured Search activation Crawling & indexing Retrieval Rerank & context Generation & citation Click & conversion Held fixed — five documents placed straight into the model’s context Measured Never observed The finding: a document already in the context gained about 41% more share of the answer when quotations were added. It is not a finding about whether your page gets retrieved — or clicked. Source: Aggarwal et al., GEO: Generative Engine Optimization, ACM KDD 2024 · aajconsult.com

There is a second detail almost nobody quotes. The five sources' shares sum to 100%, which makes the metric redistributive — one source's gain is another's loss. GEO, measured this way, is closer to a zero-sum game than a rising tide.

Google's position: it's still SEO

On 15 May 2026, Google published its first official guidance, Optimizing your website for generative AI features on Google Search, announced by John Mueller on the Search Central blog. It addresses AEO and GEO by name and says, plainly, that from Google Search's perspective optimizing for generative AI search is optimizing for the search experience — and thus still SEO.

The guide also quietly retires several tactics that GEO vendors sell: llms.txt (treated as any other text file, no special path), content chunking, special schema or Markdown versions of pages, and third-party AI-visibility tools that claim access to Google's internal ranking systems.

What actually holds up: relevance and position

A factorial experiment by Vishwakarma and colleagues — 252,000 trials across six models and eighteen factors — found that query–document relevance and position within the context are the primary determinants of which source gets cited first. The two strongest levers are being relevant to the question and being retrieved in the first place. A perfectly optimized page that never enters the model's context contributes nothing at all.

The finding that should stop you rewriting pages for AI

The SAGEO Arena study reinstated the full pipeline — crawling, retrieval, reranking, then generation — across 171,003 documents and 2,700 queries. When pages were optimized for citation in the body only, average presence in the top 20 fell by about 9%, top-10 presence after reranking fell 16%, and final citation fell 6%.

Rewriting a page to be more citable made it harder to retrieve. And since retrieval comes first, the total effect was negative.

This resolves an apparent contradiction in the field. Fixed-context studies measure what happens once you're already in the answer. End-to-end studies measure what happens across the whole pipeline. An intervention can win the first and lose the second.

Figure 2. The same rewrite gains share of the answer in fixed-context studies and loses presence across the end-to-end pipeline — because retrieval happens first.
Why GEO studies disagree: fixed-context versus end-to-end measurement Fixed-context studies place documents directly into a model's context and measure the share of the answer each one wins; rewriting for citation gained about 41 percent there. End-to-end studies restore crawling, retrieval and reranking before generation; across 171,003 documents, optimising page bodies for citation alone reduced top-20 retrieval presence by about 9 percent, top-10 presence after reranking by 16 percent, and final citation by 6 percent. The same tactic wins the first measurement and loses the second, because retrieval happens first. Why GEO studies disagree The same tactic, measured two different ways FIXED CONTEXT Five documents already in context Retrieval is skipped. Only the answer is measured. Share of the answer won +41% Quotations added · position-adjusted word count END TO END Crawl → retrieve → rerank → cite 171,003 documents. The whole pipeline runs. Effect of body-only citation rewrites Top-20 presence −9% Top-10 after rerank −16% Final citation −6% vs Both results are correct. The first measures what happens once you are already in the answer. The second measures whether you get there at all — and retrieval happens first. Sources: Aggarwal et al. KDD 2024; Kim et al. 2026 (SAGEO Arena) · aajconsult.com

C-SEO Bench tested GEO methods across roughly 1,900 queries, 16,360 documents and six domains: only three of 54 method–domain combinations were significantly positive, and none in question answering.

There is no single AI ranking

Audits consistently find that engines cite different source ecosystems entirely. 53% of domains cited by Google's AI Overviews do not appear in the organic top 10; 27% are absent from the top 100. URL-level overlap between organic Google, AI Overviews and Gemini runs at a Jaccard similarity of 0.11–0.18. There is no global "AI ranking" to optimize toward.

Being known is not being recommended

This is the finding that matters most for a startup. A 2026 study tested 112 startups across two engines. When products were named directly, ChatGPT recognized 99.4% of them. When the same products were sought through generic category discovery queries — the way an actual buyer searches before they know you exist — ChatGPT surfaced them in 3.32% of cases. Perplexity fell from 94.3% recognition to 8.29% discovery.

Figure 3. Both engines recognise nearly every startup when asked by name. Almost none surface in generic category discovery queries.
Being known is not being recommended A 2026 study tested 112 startups across two AI engines. When products were named directly, ChatGPT recognised 99.4 percent of them and Perplexity 94.3 percent. When the same products were sought through generic category discovery queries, the way a buyer searches before they know a brand exists, ChatGPT surfaced them in only 3.32 percent of cases and Perplexity in 8.29 percent. Being known is not being recommended 112 startups, asked two different ways Asked by name “Do you know [Brand]?” Asked by category “Best tools for [job]?” ChatGPT 99.4% 3.32% Perplexity 94.3% 8.29% Recognition collapses when the buyer doesn’t already know your name Tracking “does ChatGPT know us?” measures the left column. Your buyers are searching the right one. Source: Sharma, 2026 (preprint, two engines) · aajconsult.com

The model knowing your company exists tells you almost nothing about whether it will recommend you to someone who doesn't. Tracking "does ChatGPT know us" while your buyers search "best tools for X" measures the wrong thing entirely.

Measuring it without fooling yourself

Answer engines are non-deterministic. Day-to-day overlap in cited sources across four engines runs at Jaccard 0.34–0.42. Researchers suggest seven to eight repetitions per prompt as a starting point. In one audit configuration, 57.8% of ChatGPT repetitions never activated web search at all — dashboards that report "share of citations" only among answers that contained citations quietly exclude most of the sample.

Practical consequences: repeat every prompt several times, freeze your prompt set between runs, report per engine, keep presence and citation as separate numbers, count the runs where nothing was retrieved, and treat small movements as noise.

What the evidence actually supports doing

  1. Be retrievable first. Crawlable, indexed, present in the raw HTML without JavaScript. The precondition for everything else.
  2. Answer real questions directly. Relevance is the single most reproducible lever in the literature.
  3. Include verifiable, extractable evidence — dates, prices, definitions, specific figures with sources.
  4. Don't keyword stuff. It measurably backfired in the original benchmark.
  5. Don't rewrite whole page bodies purely for citability. End-to-end evidence says that can cost more in retrieval than it gains in citation.
  6. Measure per engine, with repetitions, against a frozen prompt set. Distinguish presence from citation, and category discovery from brand recognition.
  7. Publish an llms.txt if you like — it takes ten minutes and Google says it gets no special treatment. Don't call it a strategy.

The honest summary: strong evidence that already-retrieved content can be rewritten to claim more of an answer; moderate evidence that extractable, well-structured evidence helps; and in the reviewed literature, no technique demonstrating a stable, cross-platform causal effect on organic discoverability or on downstream clicks and conversions. For the craft side of using AI in your marketing without flattening your voice, see Practical Ways to Use AI in Your Marketing Without Losing Your Voice.