What Is GEO? What the Evidence Actually Says About AI Citations
By Saroj Jha · July 28, 2026 · 12 min read
Generative engine optimization is the practice of structuring content so AI answer engines — ChatGPT, Perplexity, Google's AI Overviews and AI Mode — cite it inside the answers they generate. The goal shifts from occupying a position in a list of links to being named inside a synthesized answer.
Recognition vs discovery
3.32%
discovery, against 99.4% brand recognition
Being known is not being recommended.
If you have read anything about GEO in the last two years, you have met the number: optimizing for AI increases visibility by up to 40%. It comes from a real, peer-reviewed paper. It also measured something considerably narrower than the way it gets quoted — and a July 2026 survey reviewing 45 studies lists the general version of that claim in its table of findings under a single word: rejected.
This post is what the evidence actually supports, what it doesn't, and what a Seed–Series B team should do on Monday.
Where the "40% more visibility" number came from
The foundational work is GEO: Generative Engine Optimization by Aggarwal and colleagues, published at ACM KDD 2024. It's a good paper. It named the problem and built the first benchmark — 10,000 queries, with nine content strategies tested against a control.
Here is the experimental setup, which is the part that matters. For each query, the top five Google results were placed into the model's context, one of them was rewritten under a given strategy, and the answer was measured for how much attributed text each source received.
Adding quotations moved the position-adjusted word count from 19.3 to 27.2 — a relative gain of about 41%. Adding statistics and citing sources produced smaller gains. Keyword stuffing made things worse, dropping the metric to 17.7.
So the finding is real, and it is this: a document already sitting in the model's context can be rewritten to claim a larger share of the answer. It is not a finding about whether your page gets retrieved in the first place. No clicks, referrals or conversions were observed anywhere in the study.
There is a second detail almost nobody quotes. The five sources' shares sum to 100%, which makes the metric redistributive — one source's gain is another's loss. Under the cite-sources strategy, the fifth-ranked source gained 115% while the first-ranked lost 30%. GEO, measured this way, is closer to a zero-sum game than a rising tide.
Google's position: it's still SEO
On 15 May 2026, Google published its first official guidance on the subject, Optimizing your website for generative AI features on Google Search, announced by John Mueller on the Search Central blog. It addresses AEO and GEO by name and says, plainly, that from Google Search's perspective optimizing for generative AI search is optimizing for the search experience — and thus still SEO.
The guide also quietly retires several tactics that GEO vendors sell:
- llms.txt — Google's crawler may discover the file, but treats it like any other text file. No special treatment, no preferred indexing path.
- Content chunking — no need to pre-fragment articles. Google's systems extract the relevant passage from multi-topic pages.
- Special schema or Markdown versions of pages — not required for inclusion in generative AI features.
- Third-party AI-visibility tools — Google notes that no third-party tool has access to its internal ranking or AI systems.
None of this makes structured data worthless — other engines parse it, and it costs little. It does mean that if someone is selling you an AI-visibility programme built on llms.txt and schema, the vendor with the most direct knowledge of the system has publicly said those aren't the lever.
What actually holds up: relevance and position
Strip away the tactics and two factors dominate the literature.
A factorial experiment by Vishwakarma and colleagues — 252,000 trials across six models and eighteen factors — found that query–document relevance and position within the context are the primary determinants of which source gets cited first. Separate work found that moving a source higher in the context outperforms most rewrites of the source itself.
Read that again, because it reorders the whole to-do list. The two strongest levers are being relevant to the question and being retrieved in the first place. A perfectly optimized page that never enters the model's context contributes nothing at all.
Which makes the unglamorous work decisive: can an engine crawl your page, does the content exist in the raw HTML without JavaScript executing, is the page actually indexed, and does it directly answer a question someone actually asks? Most sites that "don't show up in AI answers" fail at one of those four before any content tactic becomes relevant.
The finding that should stop you rewriting pages for AI
Here is the most useful result of the last year, and the one least likely to appear in a GEO sales deck.
The SAGEO Arena study (Kim and colleagues, peer-reviewed at ACM KDD 2026) reinstated the full pipeline — crawling, retrieval, reranking, then generation — across 171,003 documents and 2,700 queries. When pages were optimized for citation in the body only, average presence in the top 20 fell by about 9%, top-10 presence after reranking fell 16%, and final citation fell 6%.
Rewriting a page to be more citable made it harder to retrieve. And since retrieval comes first, the total effect was negative.
This resolves an apparent contradiction in the field. Fixed-context studies measure what happens once you're already in the answer. End-to-end studies measure what happens across the whole pipeline. An intervention can win the first and lose the second.
The generalizability evidence points the same way. C-SEO Bench (Puerto and colleagues, peer-reviewed in the NeurIPS 2025 Datasets & Benchmarks proceedings) tested nine GEO methods across two tasks and six domains: only three of 54 method–domain combinations were significantly positive, and none in question answering. Several transformations reduced rank. Gains also shrank as more competitors adopted the same tactics — which is what you would expect from a redistributive metric.
There is no single AI ranking
Audits consistently find that engines cite different source ecosystems entirely.
- Nearly 30% of the domains cited by Google's AI Overviews do not appear in the first-page results shown alongside them — a selection mechanism distinct from Google's own ranking (Xu, Iqbal & Montgomery, arXiv preprint 2605.14021, May 2026 — not peer reviewed; 55,393 queries over 40 days).
- Across 11,500 real user queries, the sources retrieved by organic Google, AI Overviews and Gemini overlapped at an average Jaccard similarity below 0.2 (Grossman et al., accepted for ACM SIGIR 2026; arXiv preprint 2604.27790).
There is no global "AI ranking" to optimize toward. Visibility is indexed by engine, surface, query phrasing and date. Any dashboard that blends engines into one score is averaging away the only thing you could act on.
Being known is not being recommended
This is the finding that matters most for a startup, and it deserves its own heading.
A 2026 preprint — The Discovery Gap by Amit Prakash Sharma (arXiv preprint 2601.00912, January 2026; not peer reviewed) — tested 112 Product Hunt startups across 2,240 queries to two engines. When products were named directly, ChatGPT recognized 99.4% of them. When the same products were sought through generic category discovery queries — the way an actual buyer searches before they know you exist — ChatGPT surfaced them in 3.32% of cases. Perplexity fell from 94.3% recognition to 8.29% discovery.
The model knowing your company exists tells you almost nothing about whether it will recommend you to someone who doesn't. This is a single preprint using two models, so hold it loosely — but the distinction it draws is fundamental, and it's the distinction most brand-monitoring dashboards quietly ignore. Tracking "does ChatGPT know us" while your buyers search "best tools for X" measures the wrong thing entirely.
Measuring it without fooling yourself
Answer engines are non-deterministic, and the measurement problem is genuinely harder than the optimization problem.
Tracking four engines across four industries daily for a 45-day window, Don't Measure Once (Schulte, Bleeker & Kaufmann, University of St. Gallen; arXiv preprint 2604.07585, April 2026 — not peer reviewed) found day-to-day overlap in cited sources at a Jaccard of roughly 0.34–0.42 — most of what you'd see on any given day differs from the day before. Their conclusion is that visibility is a distribution, so a single reading is one draw from a range, and every prompt needs repeating.
Then there's the denominator problem. In a large share of runs, an engine never activates web search at all — the percentages circulating for that share come from single audit configurations no one has published a method for, so we put no number on it. A dashboard that calculates your "share of citations" only among answers that contained citations, and reports that as visibility, is quietly excluding those runs.
Query phrasing changes everything too. Across 55,393 trending queries over 40 days, AI Overviews activated on 13.7% overall — but 64.7% of queries phrased as questions (Xu, Iqbal & Montgomery, arXiv preprint 2605.14021, May 2026 — not peer reviewed).
And a citation is not an endorsement. Stanford's human audit of four generative search engines (Liu, Zhang & Liang, 2023) found only 51.5% of generated sentences were fully supported by their cited sources. You can be cited for a claim your page doesn't make.
Practical consequences: repeat every prompt several times, freeze your prompt set between runs, report per engine, keep presence and citation as separate numbers, count the runs where nothing was retrieved, and treat small movements as noise until they survive repetition.
What the evidence actually supports doing
Nothing here is exotic. That's rather the point.
- Be retrievable first. Crawlable, indexed, and present in the raw HTML without JavaScript. This is the precondition for everything else, and it's where most failures live.
- Answer real questions directly. Relevance is the single most reproducible lever in the literature. Write the page that answers the question, not the page that games the answer.
- Include verifiable, extractable evidence — dates, prices, definitions, specific figures with sources. Support is moderate and conditional, and the constraint matters: fabricated statistics can increase reuse while degrading everything that makes reuse worth having.
- Don't keyword stuff. It measurably backfired in the original benchmark and hasn't recovered since.
- Don't rewrite whole page bodies purely for citability. The end-to-end evidence says that can cost you more in retrieval than it gains in citation.
- Measure per engine, with repetitions, against a frozen prompt set. Distinguish presence from citation, and category discovery from brand recognition.
- Publish an llms.txt if you like — it takes ten minutes and Google has said it gets no special treatment. Do it after everything above, and don't call it a strategy.
The honest summary of where the field stands: there is strong evidence that already-retrieved content can be rewritten to claim more of an answer; moderate evidence that extractable, well-structured evidence helps; and, in the reviewed literature, no technique demonstrating a stable, cross-platform causal effect on organic discoverability or on downstream clicks and conversions.
That isn't an argument for ignoring AI search. Buyers are genuinely researching there, and being absent is expensive. It's an argument for treating it as good information engineering with honest measurement attached — and for being sceptical of anyone selling a number. For the craft side of this — how to use AI in your marketing without flattening your voice — see Practical Ways to Use AI in Your Marketing Without Losing Your Voice.
Get the working parts free: the AI Visibility Starter Kit — a runnable audit script that scores any page for retrievability and citation-readiness, a citation prompt set, and the decision rules.
Want it built for you? The AI Visibility Sprint is the done-with-you version: four weeks, your site, a baseline across four engines, and a re-measured readout at week 10 that says plainly whether anything moved.
Sources & Further Reading
The insights in this article draw on research and thinking from these reputable sources:
Aggarwal et al. — GEO: Generative Engine Optimization (ACM KDD 2024)
The foundational GEO benchmark paper: 10,000 queries, nine content strategies, five documents placed into context. Source of the widely-quoted '~41% lift from adding quotations' finding.
https://dl.acm.org/doi/10.1145/3637528.3671900 →
Kim et al. — SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization, Proceedings of the 32nd ACM SIGKDD Conference (KDD 2026), DOI 10.1145/3770855.3818146
Peer-reviewed conference paper. End-to-end SAGEO benchmark over a 171,003-document corpus and 2,700 queries across nine domains. Source of the finding that body-only citation rewrites reduce top-20 presence ~9%, top-10 presence after reranking 16%, and final citation 6%.
https://doi.org/10.1145/3770855.3818146 →
Puerto et al. — C-SEO Bench: Does Conversational SEO Work?, NeurIPS 2025 Datasets & Benchmarks Track proceedings
Peer-reviewed conference paper. Multi-actor benchmark of nine conversational-SEO methods across two tasks and six domains; only three of 54 method–domain combinations were significantly positive, and gains shrink as adoption rises.
https://proceedings.neurips.cc/paper_files/paper/2025/hash/27aa3aeff0f8460a7b43d30fa6c5c032-Abstract-Datasets_and_Benchmarks_Track.html →
Xu, Iqbal & Montgomery — Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact, arXiv preprint 2605.14021, 13 May 2026
Preprint — not peer reviewed. 55,393 trending queries over 40 days (13 March–21 April 2026). Source of the 13.7% overall / 64.7% question-form activation rates and the finding that nearly 30% of AIO-cited domains are absent from the co-displayed first-page results.
https://arxiv.org/abs/2605.14021 →
Grossman et al. — How Generative AI Disrupts Search, accepted for ACM SIGIR 2026; arXiv preprint 2604.27790, 30 April 2026
Accepted at SIGIR 2026 (authors' camera-ready and released code state the venue); the proceedings entry is not yet published, so the linked version is the preprint. Benchmark of 11,500 real user queries comparing organic Google, AI Overviews and Gemini; retrieved sources overlap at an average Jaccard similarity below 0.2.
https://arxiv.org/abs/2604.27790 →
Schulte, Bleeker & Kaufmann — Don't Measure Once: Measuring Visibility in AI Search (GEO), arXiv preprint 2604.07585, 8 April 2026
Preprint — not peer reviewed. University of St. Gallen authors, tracking four AI search engines across four verticals over a 45-day window; day-to-day overlap in cited sources runs at a Jaccard of 0.34–0.42.
https://arxiv.org/abs/2604.07585 →
Sharma, A. P. — The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries, arXiv preprint 2601.00912, 1 January 2026
Preprint — not peer reviewed. 112 Product Hunt startups tested across 2,240 queries to ChatGPT (gpt-4o-mini) and Perplexity (sonar): 99.4% / 94.3% recognition against 3.32% / 8.29% discovery.
https://arxiv.org/abs/2601.00912 →
Liu, Zhang & Liang (2023) — Evaluating Verifiability in Generative Search Engines, arXiv:2304.09848
Stanford human audit of Bing Chat, NeevaAI, perplexity.ai and YouChat: on average only 51.5% of generated sentences are fully supported by their citations, and 74.5% of citations support the sentence they are attached to.
https://arxiv.org/abs/2304.09848 →
Martinez — Optimizing Visibility in Generative Engines: A Critical Survey of GEO (2023–2026), arXiv preprint 2607.14035, 15 July 2026
Further reading, preprint — not peer reviewed: systematic review of 45 GEO studies, including the evidence hierarchy this post's structure follows. No figure on this page rests on the survey — each is cited to the study that reports it.
https://arxiv.org/abs/2607.14035 →
Google Search Central — Optimizing your website for generative AI features on Google Search (15 May 2026)
Google's first official guidance on GEO/AEO: treats optimizing for generative AI as an SEO goal, and states llms.txt, content chunking, and special schema versions are not required for inclusion.
https://developers.google.com/search/docs/fundamentals/ai-optimization-guide →
Google Search Central Blog — A new resource for optimizing your website for generative AI (15 May 2026)
John Mueller's announcement post accompanying the guidance — reinforces Google's position that its generative features run on core Search ranking and quality systems.
https://developers.google.com/search/blog/2026/05/a-new-resource-for-optimizing →
More in AI Search & Agent Readiness
Part of the AI Search & Agent Readiness hub - see all 13 resources on this topic.
- AI agent skill: Agent Readiness Audit
- Free tool: Agent Readiness Quick Check
- Free tool: AI Answer Share-of-Voice Tracker
- Article: How AI Agents Are Changing Marketing: What Small Businesses Must Do to Stay Discoverable in 2026
Also useful in Content, Copy & SEO.