All articles
Data & Trends

Do GEO Tactics Survive Replication? What 2026 Re-Tests Found

A September 7, 2026 arXiv preprint re-tested the three tactics every GEO guide recommends and found no positive pooled effect: quotations -0.325pp, statistics -0.276pp, cite-sources -0.793pp with a confidence interval excluding zero. The query-blind page score it tested reached only rho = 0.114 across 777 query-engine groups, while adding query relevance lifted it to roughly 0.37. Covers what the replications actually measured, why volume-controlled edits fail, the Ahrefs 75,000-brand correlation table, 80% citation volatility across 536 prompt combinations, mention-versus-recommendation gaps, and a six-step programme built on what survives.

11 min read·Updated 2026-09-12

Short answer: the three tactics every GEO guide recommends — adding statistics, adding expert quotations, and citing sources — did not produce a measurable lift when re-tested at scale in September 2026. An arXiv preprint by Elisha Bajemon and Andre-Louis Rochet (7 September 2026) applied them as volume-controlled paired edits and measured no positive pooled effect: quotations −0.325 percentage points, statistics −0.276, and cite-sources −0.793 with a confidence interval excluding zero. That does not mean GEO is dead. It means the industry has been treating context-dependent tactics as universal laws, and 2026 is the year the measurement caught up.

This page is the honest version of the evidence. If you have built a content programme on statistics, quotations and citations, most of what you are doing still helps — but probably not for the reason you were told, and almost certainly not as much as a +33% or +41% headline suggests. What follows is what the replications actually tested, what they did not test, and how to restructure a GEO programme around what survives.

The five numbers that matter from the 2026 replication work: a query-blind page score predicted citation ordering at only ρ = 0.114 across 777 query-engine groups (0.118 across 450 groups on three GPT-5.x arms) · adding query relevance lifted that to roughly 0.37 · correcting query leakage cut a fitted model from 0.5724 to 0.3625 · across 536 prompt combinations, 80% of appearance cases were inconsistent between runs · and a single check matched the majority result only 72.2% of the time.

What the Original Princeton Study Claimed

The Princeton GEO paper (Aggarwal et al., arXiv:2311.09735, KDD 2024) tested nine content interventions and reported visibility lifts of +33% for statistics, +41% for expert quotations, +28% for source citations, and a −8% penalty for keyword stuffing. Those four numbers became the founding citation of commercial GEO, quoted in nearly every vendor deck and most guides on this site included. They were measured on a purpose-built benchmark, not on live engines.

That distinction is not a technicality. A benchmark fixes the query, the candidate set and the retrieval conditions. Live AI search fixes none of them. The gap between the two is exactly where the 2026 replications went looking — and where the numbers moved.

"Three popular GEO tactics showed no positive pooled lift when re-tested with volume-controlled edits. That does not mean GEO does not work. It means the query-blind score weakly predicted citation ordering, while the three historical interventions did not transfer positively."
— Elisha Bajemon & Andre-Louis Rochet, arXiv preprint, 7 September 2026 (not yet peer reviewed)

What the September 2026 Replication Actually Tested

It tested two separate objects: a deterministic 0–100 query-blind page score built from 11 text features, and a query-conditioned predictor that adds query-source relevance. The score used entropy, information and entity density, semantic coherence, self-containment, citation F1, statistic density, MMR, nDCG, redundancy and quotable density. The outcome measured was a visibility share in generated answers, not Google rank, traffic, leads or revenue.

ClaimWhat was testedResult
Page-only GEO score predicts citationsQuery-blind score vs within-query source visibilityρ = 0.114 (n = 777)
Add quotationsPaired, volume-controlled quotation edit−0.325 pp, 95% CI −0.841 to +0.171 (n = 1,531)
Add statisticsPaired, volume-controlled statistics edit−0.276 pp, 95% CI −0.970 to +0.425 (n = 1,087)
Add source citationsPaired, volume-controlled cite-sources edit−0.793 pp, 95% CI −1.533 to −0.138 (n = 1,087)
Query context mattersQuery-aware models on query-disjoint foldsρ = 0.3696 (query-only)

Source: Bajemon & Rochet, arXiv preprint, 7 September 2026. Ranking data contain 229 unique queries and 1,269 query-engine groups with five candidate sources per group. This is a preprint and has not been peer reviewed.

Why the Tactics Failed — and Why That Is Not the Whole Story

The replications edited volume, not relevance. A statistic pasted into a passage that was never about that statistic does not improve query fit; it adds tokens. The study design makes this explicit: the edits were paired and volume-controlled, applied to existing pages, which is a very different intervention from commissioning original research on the question your buyers actually ask.

The confidence intervals matter here, and most summaries of this paper skip them. The quotation interval ran from −0.841 to +0.171 and the statistics interval from −0.970 to +0.425. Both include zero, so the correct reading is "no detected effect," not "a proven negative effect." Only cite-sources produced an interval that excluded zero on the negative side.

There is also a scale question nobody has resolved. The paper measured visibility share among five candidate sources for the same query. It did not measure rankings, traffic, leads or revenue — the layers that decide whether a content programme is worth funding. Weak ordering correlation among a fixed candidate set is not the same as zero commercial effect across a portfolio of pages.

What Actually Predicts AI Citations in 2026

Query-context fit, off-page entity presence, and freshness — in that order. The preprint's own ablation is the cleanest evidence: adding query-source relevance lifted correlation from 0.114 to roughly 0.37, and correcting query leakage deflated a fitted model from 0.5724 to 0.3625. Anything that looks like a big win without query conditioning should be treated as leakage until proven otherwise.

The off-page evidence points the same direction. Ahrefs' Brand Radar study of 75,000 brands (3 September 2026) correlated every major marketing signal against how often ChatGPT, Google AI Mode and AI Overviews mention a brand:

SignalCorrelation with AI visibility
YouTube mentions0.737
Web brand mentions (unlinked)0.656 – 0.709
Anchor-text links0.511 – 0.628
Domain Rating0.266 – 0.326
Raw backlink countClose to zero

The practical translation: being named in more places beats being optimised harder. A page-level formatting edit is a small lever; getting your brand, product and people discussed across YouTube, podcasts, newsletters and comparison content is a large one. The 2026 evidence consistently ranks off-page entity presence above on-page tactics.

The Measurement Problem Nobody Wants to Talk About

Most GEO reporting is noise, because most GEO tooling samples once. Across 536 prompt-and-engine combinations with at least five checks, 80% of appearance cases were inconsistent, and a single check matched the majority result only 72.2% of the time. The working rule from practitioners is 30 to 40 runs per prompt before volatility flattens; SignalForge runs 40 runs per prompt for under 2% margin of error.

The second reporting failure is conflating mention with recommendation. Latent Space's Frontier AI visibility tracker, covering 6,762 answers across 161 categories, found Kysely appearing in 42 of 42 answers but named first choice exactly once. Category leaders also differed between GPT-5.6 Sol and GPT-6 Astra in 33 of 121 comparable categories. Share of voice built on mentions will systematically overstate commercial position.

"If AI is already talking buyers out of purchases, does your GEO reporting still stop at whether you were mentioned?"
— Paris Childress, founder of Hop AI and co-founder of GEOforge, The GEO Show, 8 September 2026

That question is now empirically grounded. A Semrush and Exploding Topics survey of 2,338 US adults (7 September 2026) found 57.5% of AI users said chatbot information caused them not to buy, while 65% said AI had at least partially replaced Google for product research. Among weekly AI users, 73.6% bought from an organic AI recommendation. AI answers both create and destroy demand, and a dashboard that only counts presence cannot see the difference.

A GEO Programme Built on What Survives

Six changes, ordered by the strength of the evidence behind them. None of them require abandoning the tactics that failed to replicate — they require repositioning those tactics from growth levers to hygiene.

  1. 1.
    Sample 30 to 40 runs per prompt before reporting anything

    With 80% of appearance cases inconsistent across five checks, single-run dashboards are measuring volatility, not visibility. Budget the runs, or report ranges instead of point estimates.

  2. 2.
    Separate appearance, citation, recommendation and sentiment

    A brand can be mentioned in every answer and chosen in none. Track the recommendation role and the sentiment of the mention separately from raw presence, then align reporting to pipeline rather than share of voice.

  3. 3.
    Optimise for query fit, not page score

    Query conditioning roughly tripled the correlation in the preprint (0.114 to ~0.37). Write each section to answer one specific sub-query, and verify it against the actual prompt your buyer uses.

  4. 4.
    Move budget from on-page formatting to off-page entity presence

    YouTube mentions correlate 0.737 and unlinked web mentions 0.656 to 0.709, against roughly zero for backlink count. Podcasts, comparison content, review sites and community threads are now higher-yield than another round of on-page editing.

  5. 5.
    Keep statistics, quotations and citations — as credibility, not as lift

    They did not replicate as visibility levers, but they remain how a passage earns trust with a human reader and signals sourcing to a retriever. Just stop forecasting +33% from them.

  6. 6.
    Track each engine separately

    Intender found only 8% source overlap across five engines and 71% of domains appearing on one engine alone. A blended share of voice hides more than it reveals.

Where We Expect the Evidence to Move Next

Three open questions will decide whether the 2026 replications hold. First, whether the interventions behave differently when the added statistics and quotations are genuinely query-relevant rather than volume-matched — the preprint lists this as unknown. Second, whether the results transfer to live consumer surfaces rather than API output, since UI and API results are known to diverge. Third, whether model-version dependence swamps content effects entirely: category leaders already differ between GPT-5.6 Sol and GPT-6 Astra in 33 of 121 categories, and Google began rolling Gemini 3.8 Flash into AI Mode in September 2026.

Until those are settled, the defensible position is uncomfortable but simple: keep doing the tactics, stop selling them as percentages.

Frequently Asked Questions

Do GEO tactics actually work?

The evidence is mixed and depends on what you mean by work. The Princeton GEO study (KDD 2024) measured +33% visibility for statistics, +41% for expert quotations and +28% for source citations. A September 7, 2026 arXiv preprint by Bajemon and Rochet re-tested those three interventions with volume-controlled paired edits and found no positive pooled effect: quotations -0.325 percentage points, statistics -0.276, and cite-sources -0.793 with a confidence interval excluding zero. The honest reading is that these tactics are context-dependent, not universal laws.

What did the September 2026 replication test actually find?

Bajemon and Rochet built a deterministic 0-100 query-blind GEO score from 11 text features including entropy, entity density, semantic coherence and statistic density. Across 777 evaluable query-engine groups it reached a pooled within-query Spearman correlation of only 0.114, replicated at 0.118 across 450 groups on three GPT-5.x arms. A query-conditioned model that adds query-source relevance reached roughly 0.37 to 0.38, which shows query context carries more of the signal than page formatting.

Why did adding statistics and quotations fail to help in the replication?

Because the test edited volume, not relevance. The three interventions were applied as paired, volume-controlled edits to existing pages, so a statistic added to a passage that was not originally about that statistic does not improve query fit. The study itself notes the uncertainty: the quotation confidence interval ran from -0.841 to +0.171 and the statistics interval from -0.970 to +0.425, both including zero. Only cite-sources produced a negative interval that excluded zero.

Is the Princeton GEO study now discredited?

No. The replication is an arXiv preprint that has not been peer reviewed, and it tested a narrower claim than the original paper. It measured whether a query-blind page score and three mechanical edits predict citation ordering among five candidate sources for the same query. That is not the same as whether a content programme built on original data and expert sourcing earns citations over time. Weak ordering correlation is not proof of zero effect.

What predicts AI citations better than page formatting?

Query-context fit. In the same preprint, adding query-source relevance lifted the correlation from 0.114 to roughly 0.3696 on query-disjoint folds, and correcting query leakage cut a fitted-model result from 0.5724 to 0.3625. Off-page evidence points the same way: the Ahrefs 75,000-brand study (September 2026) found YouTube mentions correlate 0.737 with AI visibility while raw backlink volume correlates near zero.

How volatile are AI citations between runs?

Extremely. Across 536 prompt-and-engine combinations with at least five checks, 80% of appearance cases were inconsistent, and a single check matched the majority result only 72.2% of the time on inconsistent sets. The practical rule is to run each prompt roughly 30 to 40 times before volatility flattens. SignalForge runs 40 times per prompt for under 2% margin of error, which is the same order of magnitude.

Does being mentioned by an AI engine mean being recommended?

No, and the gap is the most underreported metric in GEO. Latent Space Frontier AI visibility tracker, covering 6,762 answers across 161 categories, found Kysely appearing in 42 of 42 answers but named as first choice only once. Category leaders also differed between GPT-5.6 Sol and GPT-6 Astra in 33 of 121 comparable categories. Mention share of voice systematically overstates commercial visibility, so recommendation role and sentiment need separate tracking.

Do AI engines cite the same sources as each other?

Barely. Intender analysed 27,924 citations from 3,600 searches and found only 8% overlap across five engines, with 71% of the 4,625 cited domains appearing on exactly one engine and just 65 domains appearing in all five. SEO Pulse tested 1,500 commercial prompts and found the same top vendor selected in under 1.5% of queries. Share of voice should never be blended across engines.

Can AI answers reduce purchases rather than drive them?

Yes. A Semrush and Exploding Topics survey of 2,338 US adults published September 7, 2026 found 57.5% of AI users said chatbot information caused them not to buy, while 65% said AI had at least partially replaced Google for product research. Among weekly AI users, 73.6% bought from an organic AI recommendation. GEO reporting therefore needs to cover purchase exclusion and reputation risk, not only discovery.

What should a GEO programme measure in 2026?

Four layers, sampled repeatedly: appearance rate, whether the AI surface appeared at all; citation rate, whether your URL reached the source list; recommendation role and sentiment, whether the engine actually chose you; and downstream qualified visits and pipeline. Brainlabs measured 54 client accounts across 19 sectors and found organic sessions down 10.5% while AI referrals rose 163% and AI-driven key events rose 335%, so sessions alone badly understate the channel.

References: Bajemon, E. & Rochet, A-L. arXiv preprint, 7 September 2026 — query-blind GEO score ρ = 0.114 (n = 777), GPT-5.x replication ρ = 0.118 (n = 450), quotation −0.325 pp (95% CI −0.841 to +0.171, n = 1,531), statistics −0.276 pp (95% CI −0.970 to +0.425, n = 1,087), cite-sources −0.793 pp (95% CI −1.533 to −0.138, n = 1,087), query-conditioned ρ = 0.3696. · Aggarwal, P., Dugan, L., et al. "GEO: Generative Engine Optimization." arXiv:2311.09735, KDD 2024. · Ahrefs Brand Radar, 75,000-brand correlation study, 3 September 2026 — YouTube mentions 0.737, unlinked web mentions 0.656–0.709, anchor-text links 0.511–0.628, Domain Rating 0.266–0.326, backlinks near zero. · Intender, 27,924 citations from 3,600 searches — 8% overlap across five engines, 71% of 4,625 domains on one engine only. · Semrush & Exploding Topics survey, 2,338 US adults, 7 September 2026 — 57.5% said AI caused them not to buy; 65% replaced Google for product research; 73.6% of weekly users bought from an organic AI recommendation. · Latent Space, Frontier AI visibility tracker — 6,762 answers, 161 categories; Kysely in 42 of 42 answers, first choice once; leaders differ in 33 of 121 categories between GPT-5.6 Sol and GPT-6 Astra. · Brainlabs, Google Analytics analysis of 54 clients across 19 sectors — organic sessions −10.5%, AI referrals +163%, AI-driven key events +335%. · Trelner Research, 2 September 2026 — 7,534 Perplexity citations, 59.8% outside Tranco top 100,000. · Paris Childress, The GEO Show, 8 September 2026. · Search Engine Land — Gemini 3.8 Flash rolling into Google AI Mode, September 2026.

Want to check your site's GEO readiness?

Run the 27-point GEO audit

Related articles

AI Citation Overlap 2026: Why 67% of Brands Are Named but Only 10% of URLs Are Reused

Across 596,723 prompts answered by two or more engines, only 10.2% of cited URLs appeared on more than one engine — but 67.4% of brand names in the answer text did (Wellows, September 2026). A separate analysis of 1,851 cited sources on buying prompts found just 2.8% were brand-owned pages, and a recommended brand's own page was cited only 31% of the time. This guide explains why being named and being cited are now separate contests, and gives the 4-step playbook for competing in both.

GEO Statistics: 40+ Latest Data Points for 2025–2026

AI search traffic grew 16× from 2024 to 2026. Google dropped below 90% market share. 50% of B2B buyers start in AI chatbots. AI search converts 23× better. The most comprehensive GEO statistics collection with 46 data points, refreshed July 2026 with the latest source-attributed signals (Nico Digital, Instantpress, Axis Intelligence): Gemini ~900M MAU, AIO 48–50% coverage, ChatGPT 1B+ queries/day, 2.4T AIO citations, third-party 6.5× citation advantage.

AI Search Market Share & Growth Trends (2026 Report)

ChatGPT holds 74.78% of AI traffic and 1.2B+ monthly active users (June 2026). Gemini overtook Perplexity as #2 with 18–21.5% chatbot share and ~900M MAU (July 2026). Claude surged 320% YoY. Google dropped below 90% market share. Google AI Mode: 1B+ MAU; AI Overviews reach ~48–50% of US queries (mid-2026). The complete AI search market landscape with 2026 H1 data, refreshed July 2026 with Nico Digital and Instantpress figures.