← All articles
Fundamentals

GEO Evidence Audit 2026: What 45 Studies Actually Prove

Updated 30 September 2026 with the baseline number the field was missing: AthenaHQ's State of AI Search 2026 (29 September 2026), drawn from millions of responses across eight leading LLMs, found the average brand is mentioned in just 16.3% of AI answers about its category against 56.5% for category leaders - about 3.5 times higher - while a brand's own domain goes uncited in 84% of AI answers, Reddit supplies 21.9% of off-page citations, 37.5% of on-page crawler paths start at the blog, and Grok cites 27 domains per response against Gemini's 5. Also new: Counter-GEO-Bench (arXiv, 2 Sep 2026, EMNLP 2026) shows a brand can be cited beside misinformation it never published, and tested safety filters cut attack success by only 5.7%. Plus the July 2026 critical survey of 45 GEO studies finding no stable cross-platform causal effect, the +41% quotation lift decoded as position-adjusted word count (19.3 to 27.2), Kumar et al.'s 3-tier brand ladder (73% / 44% / 11%), and Pew Research's finding that only 1% of AI-summary visits click a cited source.

12 min read·Updated 2026-09-30

Generative Engine Optimization has a measurement problem, not a hack problem. Every week a new agency claims it "got a client cited in ChatGPT" through some formatting trick. A July 2026 critical survey reviewed 45 GEO studies and found the tactical playbook mostly does not hold up under scrutiny. This guide separates what the research actually proves from what it does not — so you can invest in the levers with reproducible evidence, not the ones that sell workshops.

The two studies that matter most for 2026 practitioners are the original Princeton GEO benchmark (Aggarwal et al., KDD 2024) and the new large-scale benchmark from Kumar et al. (arXiv:2606.20065, June 2026). Layered on top of both is the critical survey of 45 studies (Martinez, arXiv:2607.14035, July 2026) that reframes the whole field. Read together, they tell a clearer story than any single vendor case study.

What the 2026 evidence says: The Princeton quotation tactic lifted position-adjusted word count from 19.3 to 27.2 (+41%) — a share-of-answer-text gain, not clicks. The July 2026 survey of 45 studies found no technique with a stable cross-platform causal effect on organic discoverability. Kumar et al. benchmarked 100,000+ prompt responses across 100+ brands and found a 3-tier brand ladder: household names cited in 73% of answers, mid-market 44%, niche 11%. 78% of citations go to corporate websites; ranked "best-of" listicles are the top format at 21%. September 2026 update: Counter-GEO-Bench (arXiv, 2 Sep 2026) found standard safety filters cut misinformation-attack success by only ≤5.7%, while Pew Research measured real click-through on a cited source at just 1% across 12,593 AI-summary visits.

Tracing the famous "+40%" to its source

Almost every GEO pitch deck quotes "up to 40% more visibility." The July 2026 critical survey traces that number to its origin and shows exactly what it measures. In Aggarwal et al. (KDD 2024), the outcome variable is position-adjusted word count — the share of the generated answer attributed to your source, weighted by where it appears in the answer.

MetricBeforeAfterWhat it means
Quotation tactic (Princeton)19.327.2+41% relative share of answer text
ScopeFixed testbedFixed testbedSource already in model context
Does NOT measure——Clicks, retrieval, or discoverability

The critical survey is explicit: this "+41%" does not mean 40% more readers will click, nor that a page gains 40% in retrieval probability. It means that, in this testbed, a source already provided to the generator receives a larger position-weighted share of attributed text. The generalized claim that "GEO increases visibility by 40%" is listed among the claims the review rejects as unsupported. The effect is real; the interpretation sold around it is not.

The 2026 critical survey: 45 studies, one reframe

Martinez (arXiv:2607.14035, July 2026) published the first critical survey of generative engine optimization, reviewing 45 studies from November 2023 to July 2026. Its central reframing: visibility in generative engines is not a single ranking task but a stochastic, partially observable pipeline. It separates seven outcome variables rather than one score — activation, crawling and indexing, retrieval, reranking and context allocation, generation and citation, absorption and fidelity, and finally attention, clicks and conversion.

"Claims about GEO return on investment clearly outstrip the academic evidence. A tactic that improves one stage can be irrelevant, or harmful, at another. This is why a single 'AI visibility percentage' is a misleading KPI: it collapses seven different failure points into one number that cannot tell you which one is broken."
— Synthesis of Martinez, "Optimizing Visibility in Generative Engines: A Critical Survey of GEO" (arXiv:2607.14035, July 2026)

Two findings matter most for budget allocation. First, topical relevance and where content sits within the retrieved context are the only levers with reproducible evidence behind them. The generic formatting heuristics that fill most GEO checklists — bullet structuring, adding statistics, writing in a Q&A format — transfer poorly across platforms and frequently fail to replicate. Second, day-to-day source overlap is low: the same query returns a Jaccard similarity of just 0.34–0.42 across repeated runs, so a single before/after screenshot proves nothing without repeated sampling.

Large-scale benchmark: the 3-tier brand ladder (Kumar et al., 2026)

While the critical survey cools expectations on discoverability, Kumar et al. (arXiv:2606.20065, June 2026) provide the first large-scale empirical baseline for measuring GEO. They analyzed 100,000+ prompt responses across 100+ brands tracked between March and May 2026. The results quantify a clear brand-stature ladder and the formats that win citations:

FindingMetricImplication
Brand-stature ladder73% / 44% / 11%Household names, mid-market, and niche brands appear in that share of relevant answers — about 30 points per step
Source of citations78% corporate sitesAmong non-corporate sources, YouTube leads ahead of Reddit, editorial media, and Wikipedia
Top content format21% of citationsRanked "best-of" listicles are the single most-cited page type
Sentiment instability6.7× flip rateWhether a brand is framed positively or negatively flips 6.7× more often than whether it is mentioned at all

Source: Kumar et al., "Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines," arXiv:2606.20065 (June 2026).

September 2026: two results that reset the evidence standard

Two studies published in the first week of September 2026 push the field past "does GEO work" and toward a harder question: what exactly is a citation worth, and can it be trusted?

Counter-GEO-Bench: a rising citation count can be a liability

Counter-GEO-Bench (arXiv preprint, 2 September 2026, accepted to EMNLP 2026) tests whether safety systems can detect misinformation hidden inside ordinary-looking GEO documents. The researchers paired 247 human-verified queries with information-preserving rewrites and information-distorting rewrites — the distorted versions engineered to look like useful web content while steering a model toward a false claim. They evaluated three victim language models and measured attack success, false positives, and answer quality.

Defense evaluatedRelative reduction in attack successVerdict
Granite Guardian≤5.7%Marginal
Llama Guard 3≤5.7%Marginal
NeMo Self-Check Fact-Checking≤5.7% (one result not statistically significant)Not reliable
C-GEO Guard (proposed)47.6%Near-zero utility loss

Source: Counter-GEO-Bench, arXiv preprint 2 September 2026, accepted to EMNLP 2026. 247 human-verified queries; three victim language models.

"A brand may be cited next to misinformation without publishing any false content itself. In that scenario, a rising citation count can look like GEO success while the answer's framing damages trust. The strategic shift is from 'How often is my brand mentioned?' to 'How accurately and in what context is my brand represented?'"
— Synthesis of Counter-GEO-Bench (arXiv, 2 September 2026) and the accompanying practitioner analysis

The practical consequence: generic AI safety filters are not a substitute for an evidence layer. These systems are built to catch policy violations, toxic language and obvious safety problems — a polished but inaccurate product comparison passes as perfectly acceptable informational content. A GEO program should therefore track citation quality, claim fidelity, source freshness and sentiment, not mentions alone. This is the same defect the July 2026 critical survey identified from the other direction: a single "AI visibility percentage" collapses several different failure points into one number that cannot tell you which one is broken.

How often do people actually click an AI citation?

The click-loss numbers remain the best-evidenced statistics in the entire subject. Pew Research Center tracked 900 US adults through 68,879 Google searches in March 2025, of which 12,593 produced an AI summary. The results are sobering for anyone measuring GEO in sessions:

  • ▸ Users clicked a traditional result on 8% of visits where an AI summary appeared, against 15% where none did
  • ▸ Only 1% of visits to a page with an AI summary produced a click on a source cited inside it
  • ▸ Browsing sessions ended after 26% of pages with an AI summary, against 16% without
  • ▸ Separately, Pew reports 49% of US adults use AI chatbots, 42% use them to search for information, and 60% read AI-generated summaries in search results

Read alongside Seer Interactive's finding that cited pages still earn a +120% lift in clicks per impression, the picture resolves cleanly: being cited is worth a great deal relative to not being cited, and almost nothing in absolute click volume. Any GEO report that converts a citation count into projected sessions without stating this is overclaiming.

September 2026: authority is being redefined, not just disputed

The most consequential evidence published in the first half of September 2026 is not about tactics at all — it is about which signals predict being chosen. On 3 September 2026, Ahrefs published a correlation study across 75,000 brands, measuring how each marketing signal moves with the frequency that ChatGPT, Google AI Mode and AI Overviews mention a brand.

SignalCorrelation with AI visibilityRead
YouTube mentions0.737Strongest single predictor; channel ownership does not matter
YouTube view counts0.717Watched videos, not merely published ones
Web brand mentions (with or without links)0.656–0.709Unlinked coverage counts — podcasts, newsletters, roundups
Anchor-text links0.511–0.628Weaker than unlinked mentions — the classic hierarchy inverts
Domain Rating0.266–0.326Weak to moderate — a sharp demotion from traditional SEO
Raw backlinks and page counts~0No meaningful relationship
Ad spend0.215–0.286Paid media does not buy AI recommendation

Source: Ahrefs Brand Radar study of 75,000 brands, published 3 September 2026. Brand Radar samples AI answers from AI Overviews (308.3M monthly prompts), Gemini (31.5M), Perplexity (31.4M), ChatGPT (31.3M), Copilot (30.9M) and AI Mode (29.4M). Correlation is not causation.

A second September 2026 result points the same direction from the retrieval side. Trelner tested 380 software categories across Perplexity, Sonar and Sonar Pro and collected 7,534 citations: 59.8% came from domains outside the Tranco top 100,000. A product-demo vendor, GuideFlow, was the third most-cited domain — ahead of Gartner — and three related sites together held 215,128 generated best-category pages. The lesson is uncomfortable for anyone selling "authority": retrieval rewards machine-legible alignment with the query, and conventional authority is only one proxy for it.

"SEO authority and GEO citation authority are diverging. Audit which URLs actually ground competitor recommendations, not just who has the most backlinks."
— Synthesis of the Trelner Perplexity citation study (3 September 2026) and the Ahrefs 75,000-brand correlation study

Two further September 2026 datapoints set realistic expectations for measurement. SEO Pulse tested 1,500 commercial prompts and found different AI engines selected the exact same top vendor in under 1.5% of queries; more than 50% of answers mentioned brands, but only about 10% both mentioned and cited one. And Brainlabs, analysing Google Analytics from 54 clients across 19 sectors, recorded organic sessions down 10.5%, AI platform referrals up 163%, and AI-driven key events up 335% — with AI referrals generating key events at 1.5× the organic search rate. That dataset does not establish causation, but it is the clearest agency-side signal yet that volume and value are moving in opposite directions.

September 2026: the "invisible majority" benchmark

The evidence thread so far has been about tactics — which edits move a citation once a page is retrieved. On 29 September 2026, AthenaHQ published its State of AI Search 2026 report, built on millions of AI-generated responses across eight leading LLMs, and it answers a different and more uncomfortable question: how far is the average brand from being cited at all? Its answer is the single most useful baseline number published this year.

MetricValueWhat it means
Average brand mention rate16.3%The average brand appears in only about one in six AI answers about its own category
Category-leader mention rate56.5%About 3.5× the average — a narrower spread than the Kumar ladder, but the same shape
Own domain uncited84%Your own site is the source in only 16% of answers describing your products
Reddit share of off-page citations21.9%More than double the next source — community consensus outranks owned content
Crawler entry point: blog37.5%Where AI crawlers actually start on-page — not the homepage
Domains per response: Grok vs Gemini27 vs 5A 5.4× sourcing-breadth gap — never blend engines into one metric

Source: AthenaHQ, State of AI Search 2026, released 29 September 2026 (GlobeNewswire) — millions of AI responses across ChatGPT, Gemini, Perplexity, Claude, Grok and other engines.

The 84% uncited figure is the highest-value number here because it is the missing denominator in most GEO reporting. A brand can hold a stable share of the answers that mention it and still be invisible in five of every six answers where buyers are deciding. It also explains why the AthenaHQ finding "informational and comparative formats drive more than half of all AI citations" matters more than any formatting trick: the answers you are absent from are overwhelmingly comparison answers.

"AI search is a completely different ecosystem than traditional search... One thing AI models have in common is the tendency to return to sources that have established trust. Brands who invest in citable content now are building long-term credibility that will be difficult to dislodge down the road."
— Andrew Yan, CEO and Co-Founder, AthenaHQ (State of AI Search 2026, 29 September 2026)

Read against the rest of this audit, the September benchmarks line up into one coherent story rather than four competing ones:

QuestionStrongest 2026 evidenceAnswer
Can a page-only score predict citations?Bajemon & Rochet (arXiv, 7 Sep 2026): ρ = 0.114 query-blindWeakly — and query relevance lifts it to ~0.37
Do the original tactics replicate?Same paper: quotations −0.325 pp, statistics −0.276 ppNo positive pooled effect
What predicts being mentioned?Ahrefs 75,000 brands: YouTube mentions 0.737, backlinks ~0Earned mentions, not links
How far is the average brand from visible?AthenaHQ, 8 LLMs: own domain uncited in 84% of answersVery far, but the gap is closing slowly

What this means for your GEO program

The evidence does not say "stop doing GEO." It says scope it to the stage you can actually control. Four takeaways:

  1. 1.
    Treat indexing as the real bottleneck. The July 2026 survey is blunt: no technique yet shows a stable effect on organic discoverability. Get into every engine index first — see the get-indexed-by-AI-search-engines guide. Optimization only pays off once you are in the retrieval pool.
  2. 2.
    Optimize for citation, measured correctly. The Princeton lifts are real for already-retrieved pages: expert quotations +41%, statistics +33%, fluency +29%, citations +28%. But measure with repeated sampling across prompts, not single screenshots — day-to-day overlap is only 0.34–0.42.
  3. 3.
    Build entity authority, not just on-page tricks. Kumar et al. show a 3-tier brand ladder and that 78% of citations go to corporate sites while best-of listicles win 21% of citations. Topical relevance and early context placement are the only levers with reproducible evidence — earn them through real coverage, not keyword stuffing (which costs −8%).
  4. 4.
    Measure accuracy, not just mentions. Counter-GEO-Bench (September 2026) shows a brand can be cited alongside distorted claims it never published, and that standard safety filters cut attack success by no more than 5.7%. Track claim fidelity and sentiment alongside citation volume, and accept that Pew's data puts real click-through on a cited source at roughly 1% — treat citation as an authority signal, not a traffic forecast.

Frequently asked questions

What did the 2026 GEO critical survey find?

The July 2026 survey (Martinez, arXiv:2607.14035) reviewed 45 GEO studies from November 2023 to July 2026 and concluded that GEO techniques reliably change how an already-retrieved page is cited, but no reviewed technique shows a stable, cross-platform causal effect on organic discoverability. It is the first systematic audit of the field and a useful corrective to overclaimed tactical wins.

What does the +41% GEO lift actually measure?

It measures position-adjusted word count — the share of the AI answer attributed to a source, weighted by position — rising from 19.3 to 27.2 in Aggarwal et al. (KDD 2024). That is a real, reproducible citation-gain signal for already-retrieved pages. It is not a click gain or a retrieval gain, which is the part most pitches get wrong.

Is there large-scale evidence that GEO works?

Yes, at the citation stage. Kumar et al. (arXiv:2606.20065, June 2026) analyzed 100,000+ prompt responses across 100+ brands and found a brand-stature ladder: household names in 73% of relevant answers, mid-market 44%, niche 11%. About 78% of citations go to corporate websites, and ranked best-of listicles are the most-cited format at 21%.

Does GEO guarantee my content gets found by AI?

No. The 2026 survey explicitly separates discoverability — crawling, indexing and retrieval — from citation. Most GEO tactics only move the needle once a page is already in the engine retrieval pool. Indexing and retrieval remain the binding constraint, not formatting tricks. Start with the indexing guide.

Should I still do GEO optimization?

Yes, but scope it correctly. The Princeton lifts are real for already-retrieved pages: quotations +41%, statistics +33%, fluency +29%, citations +28%. Pair content optimization with the indexing and freshness work that actually gets you into the retrieval pool, and measure with repeated sampling rather than single screenshots.

What is the best content format for AI citations?

Kumar et al. (2026) found ranked best-of listicles are the most-cited format at 21% of all citations, ahead of every other structure. Combined with the Princeton finding that statistics (+33%) and expert quotations (+41%) lift citation, a well-sourced listicle with named data is a strong GEO template.

Can a rising AI citation count actually be a bad sign?

Yes. Counter-GEO-Bench (arXiv preprint, 2 September 2026, accepted to EMNLP 2026) demonstrated that a brand can be cited next to misinformation it never published. The tested safety filters — Granite Guardian, Llama Guard 3 and NeMo Self-Check — reduced attack success by no more than 5.7% relative, and one reduction was not statistically significant. Track claim fidelity, sentiment and source freshness alongside raw citation volume.

How often do users click a source cited in an AI summary?

Rarely. Pew Research tracked 900 US adults through 68,879 Google searches in March 2025, of which 12,593 produced an AI summary. Only 1% of visits to a page with an AI summary produced a click on a cited source, and users clicked a traditional result on 8% of AI-summary visits versus 15% where no summary appeared. Cited pages do earn about +120% more clicks per impression than uncited ones (Seer Interactive) — but absolute volume stays small.

Should small or niche brands bother with GEO?

Yes, but with realistic expectations. The 2026 large-scale benchmark shows niche brands appear in only 11% of relevant answers versus 73% for household names — about 30 points per tier. The lever with reproducible evidence is topical relevance and early context placement, which smaller brands can win on specific subtopics even without broad authority.

How often is the average brand mentioned or cited in AI answers?

Rarely. AthenaHQ's State of AI Search 2026 report, published 29 September 2026 from millions of responses across eight leading LLMs, found the average brand is mentioned in just 16.3% of AI answers about its category, while category leaders reach 56.5% — about 3.5 times the average. More striking still, a brand's own domain goes uncited in 84% of AI answers, meaning most descriptions of a product come from third-party sources rather than the company's own site. Reddit alone supplies 21.9% of off-page citations, more than double the next source.

Where do backlinks and domain authority rank?

Barely. The Ahrefs study of 75,000 brands (published 3 September 2026) found raw backlink counts and page counts correlate close to zero with how often AI systems mention a brand, and Domain Rating only 0.266–0.326. Mentions dominate: YouTube mentions 0.737, YouTube view counts 0.717, unlinked web brand mentions 0.656–0.709, anchor-text links 0.511–0.628, ad spend 0.215–0.286. Unlinked mentions beat linked ones — the inverse of the classic SEO hierarchy.

References: Aggarwal, P., Dugan, L., et al. "GEO: Generative Engine Optimization." arXiv:2311.09735, KDD 2024. · Martinez, "Optimizing Visibility in Generative Engines: A Critical Survey of GEO." arXiv:2607.14035 (July 2026) — review of 45 studies, Nov 2023–Jul 2026. · Kumar et al. "Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines." arXiv:2606.20065 (June 2026) — 100K+ prompt responses across 100+ brands. · Princeton GEO benchmark (GEO-bench, 10,000 queries × 9 datasets). · Capston AI, "The GEO Evidence Audit: What 45 Studies Actually Prove" (2026 synthesis of the critical survey). · AI Eating the World, "GEO Has a Measurement Problem, Not a Hack Problem" (2026). · Counter-GEO-Bench, arXiv preprint (2 September 2026, accepted EMNLP 2026) — 247 human-verified queries, three victim language models, C-GEO Guard 47.6% relative reduction in attack success. · Pew Research Center, AI Overviews click-behavior study (March 2025 data; 900 US adults, 68,879 searches, 12,593 with AI summaries; 1% cited-source click rate; 49% of US adults use AI chatbots). · AthenaHQ, State of AI Search 2026 (29 September 2026) — millions of AI responses across eight leading LLMs; average brand mentioned in 16.3% of answers, category leaders 56.5%, own domain uncited in 84%, Reddit 21.9% of off-page citations, blog is 37.5% of on-page crawler paths, Grok sources 27 domains per response versus Gemini's 5; quote from Andrew Yan, CEO and Co-Founder.

Want to check your site's GEO readiness?

Run the 27-point GEO audit

Related articles

What Is GEO (Generative Engine Optimization)? Complete Guide

Updated October 2026: GEO is the practice of optimizing content to be cited and referenced by AI search engines like ChatGPT Search, Perplexity, Google AI Overviews, Gemini, and Claude. New this month: entity disambiguation — the GEO gate nobody budgeted for — with the 85%-third-party-mention finding, the 6.5x third-party citation advantage, a four-layer entity stack (naming consistency, disambiguatingDescription, sameAs, Wikidata QID) and why 82% of commercial-intent citations go to third parties. Plus the September baseline: Ahrefs Brand Radar counted 462.8M prompts per month across six AI surfaces, SparkToro measured 68.01% of Google searches ending without a click, and a 7 September 2026 replication found volume-controlled quotation, statistics and citation edits produced no pooled lift - while query relevance tripled the predictive score from 0.114 to roughly 0.37. Plus the Princeton KDD 2024 strategy lifts and the 75,000-brand mention-versus-backlink inversion.

GEO vs SEO: 7 Critical Differences You Need to Know (2026 Update)

SEO targets keyword rankings and clicks. GEO targets AI citations and brand mentions. With AI search traffic growing 527% YoY in 2026, Google AIO covering 48-50% of queries with 62-83% of sources outside organic top 10, and Gartner predicting 25% search volume decline, this guide breaks down the 7 key differences with fresh 2026 data and verified statistics — and as of late August 2026, 32% of marketing leaders rank GEO their #1 2026 priority (BrightEdge).

How AI Search Engines Work: RAG Architecture Explained

Updated September 2026: Google confirmed there is no separate AI index for AI Overviews and AI Mode and told publishers to deprioritize AEO/GEO hacks such as content chunking. Three new studies qualify the classic four-stage RAG model: a replication found no positive pooled effect for quotations (-0.325 pp), statistics (-0.276 pp) or source citations (-0.793 pp); Trellner found 59.8% of Perplexity citations come from domains ranked worse than #100,000; and Intender found only 8% source overlap across five engines. Complete 4-stage pipeline breakdown with stage-by-stage optimization for ChatGPT, Perplexity, Gemini, Claude and Google AI Mode.