Beyond SEO: The Science of GEO
What peer-reviewed research and large-scale industry data actually reveal about getting cited by generative engines — including the study that says it barely works.
Read the interactive version →Key findings
- GEO is a real, peer-reviewed research field — not just a marketing buzzword. The founding study (Princeton, Georgia Tech, Allen Institute for AI, IIT Delhi) measured up to 40% visibility gains in generative-engine answers from content changes alone, across a 10,000-query benchmark. [1][2]
- What worked and what failed inverts classic SEO instinct: adding statistics, quotations and source citations lifted visibility, while old-fashioned keyword stuffing reduced it. [1][3]
- The strongest counter-evidence is also peer-reviewed: C-SEO Bench (NeurIPS 2025) tested nine conversational-SEO methods across 1,921 queries and found they 'generally do not introduce any gain' — and that gains shrink as more actors adopt the same tactics. [5]
- AI citations are decoupling from Google rankings: pages cited by AI Overviews that also rank top-10 fell from ~76% (July 2025) to ~38% (March 2026) in Ahrefs' data — partly real change, partly better measurement. [6][7]
- AI-referred traffic is small but unusually valuable: Semrush measured AI search visitors converting at ~4.4x the rate of traditional organic visitors, and for Ahrefs' own site, 0.5% of visitors from AI assistants drove 12.1% of signups. [9][10]
01A discipline is born in a lab, not an agency
In November 2023, researchers from Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi posted a paper that quietly named a new field: 'GEO: Generative Engine Optimization.' It was accepted at KDD 2024, one of the top peer-reviewed venues in data science. [1] The authors built GEO-bench — a benchmark of 10,000 queries drawn from nine data sources — and simulated a generative engine that retrieves five competing web sources and synthesizes an answer with citations. [2] Then they systematically modified one source at a time and measured how much of the final answer cited it.
The headline result: black-box content optimizations improved a source's visibility in generative-engine responses by up to 40%. The paper's own caveat travels with that number everywhere in this report — efficacy varies across domains, and the 40% is a ceiling in a controlled five-source setup, not an average expectation. [1]
02The equalizer effect
Buried in the paper is its most strategically interesting result: GEO helped lower-ranked sources most. In the reported breakdowns, sources sitting at position five of the retrieval ranking saw visibility gains above 100% from citation-style optimizations — the project page itself states that GEO 'will benefit websites lower ranked in search engines more.' [2][3]
If that generalizes, it matters: in classic SEO, incumbency compounds — the sites that rank get the clicks that reinforce the rank. In answer engines, the door is open wider for challengers whose content is evidence-dense and citable, even without a decade of domain authority. The word 'if' is doing real work in that sentence, which brings us to the counter-evidence.
03The study that says it barely works
Intellectual honesty requires leading with the strongest objection. In 2025, researchers at Parameter Lab and collaborators published C-SEO Bench, accepted at NeurIPS 2025 — a benchmark testing nine conversational-SEO methods across two tasks, six domains and 1,921 queries. Their conclusion was blunt: conversational SEO methods 'generally do not introduce any gain,' traditional SEO strategies outperform them, and whatever gains exist shrink as more actors adopt the same tactics, because visibility in an answer is zero-sum. [5]
04Whom do the engines actually cite?
Benchmarks are simulations; citation studies watch the real machines. Ahrefs analyzed 863,000 keyword SERPs containing four million AI Overview citations and found that only 37.9% of cited URLs ranked in Google's top 10 for the query — 31% ranked beyond the top 100 entirely. [6] Eight months earlier, the comparable figure in Ahrefs' own data had been roughly 76%. [7]
Reddit · Wikipedia · Youtube · Linkedin
Who wins citations is equally revealing. Semrush's three-month study of roughly 325,000 prompts across ChatGPT, Google AI Mode and Perplexity found Reddit and Wikipedia topping citation charts across engines, with LinkedIn appearing in 14.3% of ChatGPT Search responses — yet even the biggest domain rarely exceeded ~5% of total citations on a platform. [8] The citation economy is simultaneously dominated by user-generated platforms and radically long-tailed: there is room, but it is earned page by page.
05Small traffic, disproportionate value
The commercial case for caring about any of this rests on what AI-referred visitors do after they arrive. Three independent data points, each with its caveat, point the same direction.
The caveats are non-trivial: Semrush and Ahrefs sell search tooling, the Ahrefs figure is n = 1 company in B2B SaaS, and Adobe's own data showed AI-referred retail traffic converting worse than average through most of 2025 before the pattern flipped. [9][10][11][12] What survives every caveat is the mechanism: a visitor who arrives after an AI explained, compared and pre-sold the options shows up far closer to a decision.
06The money has already voted
Whatever the benchmarks conclude, capital is behaving as if GEO is real. Profound — a platform that monitors and optimizes how brands appear in AI answers — raised a $35M Series B led by Sequoia in August 2025, citing customers including Ramp, US Bank, Indeed and MongoDB, then raised a further $96M by February 2026. [13][14] Vendor surveys report the demand side moving too, with enterprise marketing leaders directing meaningful budget shares toward answer-engine optimization — though those surveys come from firms selling exactly that service, and we weight them accordingly.
It is worth remembering how fast this became orthodoxy: Gartner's February 2024 prediction that AI chatbots would cut traditional search volume 25% by 2026 [15] did not age well as a literal forecast — but as a budget-moving narrative, it may be the most commercially consequential analyst note of the decade.
07The measurement problem nobody solved yet
One finding cuts across every optimization claim in this field: generative answers are unstable. SE Ranking's research on Google's AI Mode found the engine changed at least one cited source for roughly half of keywords when the identical search was simply repeated, with very high URL turnover across a month — although the single most-cited URL stayed stable for most keywords tested. [16] Academic work confirms the deeper cause: large language models are rarely deterministic even at settings that promise determinism. [4]
Every search tactic in history has died the same death: it worked until everyone did it.
Keyword stuffing died. Link farms died. Content mills died. And the counter-study in this report predicts the same funeral for 'GEO tricks' — the moment everyone adds statistics to their pages, adding statistics stops being an edge. So is this whole field pointless? Look closer at what the Princeton paper actually rewarded. Evidence. Quotations. Named sources. It didn't discover a new trick; it discovered that the machine checks your homework. The tactic that worked is called being verifiable. That one cannot commoditize — because it isn't a tactic. It's a reputation.
Think of an AI engine as the most well-read editor who ever lived, with one flaw: it can only recommend what it can verify. It doesn't care about your ad budget. It cares whether, everywhere it looks, your story checks out. Twenty years of SEO taught brands to impress an algorithm. The next twenty will teach them something older — how to earn a reference.
That is why we call our work authority engineering, not engine optimization: engines change every quarter; the habit of trusting verifiable sources is permanent. This is our interpretation, clearly fenced from the findings — and note that we published the strongest study against our own thesis, right above. That is what being a source worth citing looks like.
Limitations
- The per-tactic percentages from the GEO paper (~+41% statistics, ~+28% quotations, ~−10% keyword stuffing) were cross-checked against secondary analyses of the paper rather than re-extracted from its tables; the abstract-level 'up to 40%' claim is verified directly.
- GEO-bench simulates a generative engine with five competing sources; real answer engines retrieve differently, and real competitors do not stand still.
- Semrush, Ahrefs, BrightEdge, SE Ranking and Conductor all sell search or AEO tooling; their studies are large but not disinterested.
- Conversion multipliers (4.4x, ~23x) come from different verticals, definitions and periods and must not be averaged or compared directly.
- Citation-pattern studies describe engines as they behaved during specific windows in 2025–2026; these systems are updated continuously and figures decay quickly.
- No study yet isolates GEO tactics from confounders like brand strength, PR coverage and link authority — correlation between citable content and citations is not proof of causation.
Conclusion
Is there a science of Generative Engine Optimization? The beginnings of one — with real benchmarks, real effect sizes, and a real replication fight. Content that carries evidence gets cited more [1]; tactics commoditize fast [5]; citations are drifting away from rankings [6]; and the traffic at stake, while small, converts like nothing else in the acquisition mix. [9][10]
The strategic bet is not on any single tactic. It is on the direction all four findings share: machines are learning to allocate attention the way careful editors always did — toward sources that can be checked. Brands that become checkable, quotable and consistently corroborated are optimizing for every engine at once, including the ones not built yet.
Sources
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande — GEO: Generative Engine Optimization (KDD 2024) (November 2023)
- GEO authors (Princeton et al.) — GEO project page & GEO-bench benchmark (2024)
- Blck Alpaca — The Princeton GEO Study: Methodology, Results and Critique (2025)
- arXiv preprint — Non-Determinism of 'Deterministic' LLM Settings (August 2024)
- Parameter Lab et al. — C-SEO Bench: Does Conversational SEO Work? (NeurIPS 2025 Datasets & Benchmarks) (June 2025)
- Ahrefs (Louise Linehan) — 38% of Pages Cited in AI Overviews Also Rank in the Top 10 (March 2026)
- Search Engine Journal — Google AI Overview Citations From Top-Ranking Pages Drop Sharply (2026)
- Semrush — The Most-Cited Domains in AI: A 3-Month Study (2026)
- Semrush — AI Search Traffic Study — AI visitors convert at 4.4x organic (June 2025)
- Ahrefs — Does AI Search Traffic Convert Better Than Traditional Search? For Ahrefs, Yes (June 2025)
- Adobe — Adobe Analytics: Traffic to US retail websites from generative AI sources jumps 1,200 percent (March 2025)
- Digital Commerce 360 (Adobe data) — Adobe: AI-referred traffic to retail sites doubles in a year, converts 54% better (June 2026)
- Profound / PR Newswire — Profound Raises $35M Series B as AI Search Becomes the Next Platform Shift (August 2025)
- Fortune — As AI threatens search, Profound raises $96 million to help brands stay visible (February 2026)
- Gartner — Gartner Predicts Search Engine Volume Will Drop 25% by 2026 (February 2024)
- SE Ranking — AI Mode Research: Sources, Volatility & Differences between AIO and Organic Search (2025)