BMR-2026-002 · GENERATIVE ENGINE OPTIMIZATION

Beyond SEO: The Science of GEO

What peer-reviewed research and large-scale industry data actually reveal about getting cited by generative engines — including the study that says it barely works.

AUGUST 202613 min read 16 sources2023–2026 GLOBALBy Berke Kızılcan — Founder & CEO, BEMACCI
Read the interactive version →
AI visibility gain
+40%
from content changes alone — 10,000-query benchmark
Source: Aggarwal et al., 'GEO: Generative Engine Optimization', KDD 2024

Key findings

01A discipline is born in a lab, not an agency

In November 2023, researchers from Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi posted a paper that quietly named a new field: 'GEO: Generative Engine Optimization.' It was accepted at KDD 2024, one of the top peer-reviewed venues in data science. [1] The authors built GEO-bench — a benchmark of 10,000 queries drawn from nine data sources — and simulated a generative engine that retrieves five competing web sources and synthesizes an answer with citations. [2] Then they systematically modified one source at a time and measured how much of the final answer cited it.

The headline result: black-box content optimizations improved a source's visibility in generative-engine responses by up to 40%. The paper's own caveat travels with that number everywhere in this report — efficacy varies across domains, and the 40% is a ceiling in a controlled five-source setup, not an average expectation. [1]

Effect of content tactics on generative-engine visibility (GEO-bench)
Adding statistics~+41%
Adding quotations~+28%
Citing sources (strongest for low-ranked sites)strong ↑
Keyword stuffing~−10%
Source: Aggarwal et al., KDD 2024, as reported in secondary analyses — exact values vary by metric
In the first peer-reviewed GEO study, which content tactic actually REDUCED a page's visibility in AI answers?
Answer: Keyword stuffing
Keyword stuffing — the defining tactic of old-school SEO — performed roughly 10% worse than doing nothing, while adding statistics, quotations and citations lifted visibility. Generative engines reward what editors reward.
Source: Aggarwal et al., KDD 2024
The inversion: The tactics that lifted AI visibility are the tactics of credible writing — evidence, attribution, quotable clarity. The tactic that defined spam-era SEO, keyword stuffing, actively hurt. Generative engines, at least in this benchmark, reward what editors reward.

02The equalizer effect

Buried in the paper is its most strategically interesting result: GEO helped lower-ranked sources most. In the reported breakdowns, sources sitting at position five of the retrieval ranking saw visibility gains above 100% from citation-style optimizations — the project page itself states that GEO 'will benefit websites lower ranked in search engines more.' [2][3]

If that generalizes, it matters: in classic SEO, incumbency compounds — the sites that rank get the clicks that reinforce the rank. In answer engines, the door is open wider for challengers whose content is evidence-dense and citable, even without a decade of domain authority. The word 'if' is doing real work in that sentence, which brings us to the counter-evidence.

03The study that says it barely works

Intellectual honesty requires leading with the strongest objection. In 2025, researchers at Parameter Lab and collaborators published C-SEO Bench, accepted at NeurIPS 2025 — a benchmark testing nine conversational-SEO methods across two tasks, six domains and 1,921 queries. Their conclusion was blunt: conversational SEO methods 'generally do not introduce any gain,' traditional SEO strategies outperform them, and whatever gains exist shrink as more actors adopt the same tactics, because visibility in an answer is zero-sum. [5]

Does GEO actually work?
Aggarwal et al. — KDD 2024
Content optimizations improved generative-engine visibility by up to 40% on a 10,000-query benchmark; statistics, quotations and citations drove the gains, with the largest effects for low-ranked sources.
Parameter Lab et al. — C-SEO Bench, NeurIPS 2025
Across 9 methods, 6 domains and 1,921 queries, conversational-SEO tactics 'generally do not introduce any gain'; traditional SEO outperforms, and effects decay toward zero as adoption spreads.
Both are credible, peer-reviewed benchmarks with different setups. The GEO paper measures a controlled five-source sandbox where one competitor optimizes and the rest stand still; C-SEO Bench tests broader tasks and models the congestion that arrives when everyone optimizes. The defensible synthesis: evidence-dense, citable content demonstrably moves the needle in isolation, and the advantage erodes as it becomes table stakes — an argument for moving early, not for moving never.

04Whom do the engines actually cite?

Benchmarks are simulations; citation studies watch the real machines. Ahrefs analyzed 863,000 keyword SERPs containing four million AI Overview citations and found that only 37.9% of cited URLs ranked in Google's top 10 for the query — 31% ranked beyond the top 100 entirely. [6] Eight months earlier, the comparable figure in Ahrefs' own data had been roughly 76%. [7]

How tightly do AI citations track rankings?
Ahrefs — July 2025
~76% of AI Overview citations ranked in the organic top 10 — suggesting AI visibility is largely a byproduct of classic SEO.
Ahrefs — March 2026
37.9% of cited URLs rank top-10; 31% rank beyond position 100 — suggesting citations are decoupling from rankings.
Ahrefs itself cautions that part of the 76→38 shift reflects improved citation-detection in its own tooling, not only changed Google behavior. The honest statement: overlap between rankings and AI citations is real but shrinking, and the measurement itself is unsettled. [[6]][[7]]
Who the AI engines cite most
Reddit · Wikipedia · Youtube · Linkedin
Across ~325,000 prompts on ChatGPT, Google AI Mode and Perplexity, user-generated platforms dominate the citation charts — yet no single domain exceeds ~5% of citations. (Semrush, 2026)

Who wins citations is equally revealing. Semrush's three-month study of roughly 325,000 prompts across ChatGPT, Google AI Mode and Perplexity found Reddit and Wikipedia topping citation charts across engines, with LinkedIn appearing in 14.3% of ChatGPT Search responses — yet even the biggest domain rarely exceeded ~5% of total citations on a platform. [8] The citation economy is simultaneously dominated by user-generated platforms and radically long-tailed: there is room, but it is earned page by page.

05Small traffic, disproportionate value

The commercial case for caring about any of this rests on what AI-referred visitors do after they arrive. Three independent data points, each with its caveat, point the same direction.

4.4x
Conversion rate of AI search visitors versus traditional organic visitors, in Semrush's traffic study.
Semrush, June 2025
12.1%
Share of Ahrefs' signups driven by AI-assistant referrals that made up just 0.5% of its visitors — a ~23x rate. Single-company data.
Ahrefs first-party data, 2025
+54%
Conversion advantage of AI-referred retail traffic over non-AI sources by mid-2026 — reversing 2025, when it converted worse.
Adobe data via Digital Commerce 360, June 2026

The caveats are non-trivial: Semrush and Ahrefs sell search tooling, the Ahrefs figure is n = 1 company in B2B SaaS, and Adobe's own data showed AI-referred retail traffic converting worse than average through most of 2025 before the pattern flipped. [9][10][11][12] What survives every caveat is the mechanism: a visitor who arrives after an AI explained, compared and pre-sold the options shows up far closer to a decision.

06The money has already voted

Whatever the benchmarks conclude, capital is behaving as if GEO is real. Profound — a platform that monitors and optimizes how brands appear in AI answers — raised a $35M Series B led by Sequoia in August 2025, citing customers including Ramp, US Bank, Indeed and MongoDB, then raised a further $96M by February 2026. [13][14] Vendor surveys report the demand side moving too, with enterprise marketing leaders directing meaningful budget shares toward answer-engine optimization — though those surveys come from firms selling exactly that service, and we weight them accordingly.

It is worth remembering how fast this became orthodoxy: Gartner's February 2024 prediction that AI chatbots would cut traditional search volume 25% by 2026 [15] did not age well as a literal forecast — but as a budget-moving narrative, it may be the most commercially consequential analyst note of the decade.

07The measurement problem nobody solved yet

One finding cuts across every optimization claim in this field: generative answers are unstable. SE Ranking's research on Google's AI Mode found the engine changed at least one cited source for roughly half of keywords when the identical search was simply repeated, with very high URL turnover across a month — although the single most-cited URL stayed stable for most keywords tested. [16] Academic work confirms the deeper cause: large language models are rarely deterministic even at settings that promise determinism. [4]

Practical consequence: Any vendor claiming to measure your 'AI visibility' from a handful of prompt runs is selling noise. Credible measurement requires repeated sampling over time — and credible optimization targets the stable core of citations, not the churning edge.
BEMACCI Perspective — editorial commentary, not research findings

Every search tactic in history has died the same death: it worked until everyone did it.

Keyword stuffing died. Link farms died. Content mills died. And the counter-study in this report predicts the same funeral for 'GEO tricks' — the moment everyone adds statistics to their pages, adding statistics stops being an edge. So is this whole field pointless? Look closer at what the Princeton paper actually rewarded. Evidence. Quotations. Named sources. It didn't discover a new trick; it discovered that the machine checks your homework. The tactic that worked is called being verifiable. That one cannot commoditize — because it isn't a tactic. It's a reputation.

Think of an AI engine as the most well-read editor who ever lived, with one flaw: it can only recommend what it can verify. It doesn't care about your ad budget. It cares whether, everywhere it looks, your story checks out. Twenty years of SEO taught brands to impress an algorithm. The next twenty will teach them something older — how to earn a reference.

That is why we call our work authority engineering, not engine optimization: engines change every quarter; the habit of trusting verifiable sources is permanent. This is our interpretation, clearly fenced from the findings — and note that we published the strongest study against our own thesis, right above. That is what being a source worth citing looks like.

Limitations

Conclusion

Is there a science of Generative Engine Optimization? The beginnings of one — with real benchmarks, real effect sizes, and a real replication fight. Content that carries evidence gets cited more [1]; tactics commoditize fast [5]; citations are drifting away from rankings [6]; and the traffic at stake, while small, converts like nothing else in the acquisition mix. [9][10]

The strategic bet is not on any single tactic. It is on the direction all four findings share: machines are learning to allocate attention the way careful editors always did — toward sources that can be checked. Brands that become checkable, quotable and consistently corroborated are optimizing for every engine at once, including the ones not built yet.

Sources

  1. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande — GEO: Generative Engine Optimization (KDD 2024) (November 2023)
  2. GEO authors (Princeton et al.) — GEO project page & GEO-bench benchmark (2024)
  3. Blck Alpaca — The Princeton GEO Study: Methodology, Results and Critique (2025)
  4. arXiv preprint — Non-Determinism of 'Deterministic' LLM Settings (August 2024)
  5. Parameter Lab et al. — C-SEO Bench: Does Conversational SEO Work? (NeurIPS 2025 Datasets & Benchmarks) (June 2025)
  6. Ahrefs (Louise Linehan) — 38% of Pages Cited in AI Overviews Also Rank in the Top 10 (March 2026)
  7. Search Engine Journal — Google AI Overview Citations From Top-Ranking Pages Drop Sharply (2026)
  8. Semrush — The Most-Cited Domains in AI: A 3-Month Study (2026)
  9. Semrush — AI Search Traffic Study — AI visitors convert at 4.4x organic (June 2025)
  10. Ahrefs — Does AI Search Traffic Convert Better Than Traditional Search? For Ahrefs, Yes (June 2025)
  11. Adobe — Adobe Analytics: Traffic to US retail websites from generative AI sources jumps 1,200 percent (March 2025)
  12. Digital Commerce 360 (Adobe data) — Adobe: AI-referred traffic to retail sites doubles in a year, converts 54% better (June 2026)
  13. Profound / PR Newswire — Profound Raises $35M Series B as AI Search Becomes the Next Platform Shift (August 2025)
  14. Fortune — As AI threatens search, Profound raises $96 million to help brands stay visible (February 2026)
  15. Gartner — Gartner Predicts Search Engine Volume Will Drop 25% by 2026 (February 2024)
  16. SE Ranking — AI Mode Research: Sources, Volatility & Differences between AIO and Organic Search (2025)
bemacci.com

See the whole picture

This page is one detail view. The work, the system, the client roster and the founder story all live on the main site — with the full experience behind them.

Open bemacci.com Start a project