All articles
Fundamentals

The Princeton GEO Study: Benchmark & Findings Explained

The Princeton/IIT Delhi/Georgia Tech GEO paper (KDD 2024) tested 9 optimization strategies on 10,000 queries. Updated September 2026 with new validations: Seer Interactive quantified the click value of a citation at +120% for cited brands, Ahrefs confirmed only about 38% of AI Overview citations come from organic positions 1-10, and the August 2026 Reddit collapse (-86.4% of ChatGPT citations in four days) shows exactly where page-level optimization stops working. Added September 2026: the AgentGEO failure taxonomy attributes 62.2% of citation failures to semantic alignment and only 27.1% to the content-quality bucket the nine tactics target, and a 12,500-query ConvertMate benchmark found 83% of AI Overview citations go to pages outside the organic top 10.

18 min read·Updated 2026-09-07

The paper "GEO: Generative Engine Optimization" (Aggarwal, Dugan, et al., KDD 2024) is the first systematic, benchmarked study of how content modifications affect citation visibility in AI-generated answers. It introduced GEO-bench — a benchmark of 10,000 real search queries across 9 datasets — and tested 9 optimization strategies with quantified visibility results. As of September 3, 2026, it remains the most methodologically rigorous reference study in the GEO field, with directional findings validated by Conductor, GrackerAI, BrightEdge, and Axis Intelligence.

The paper was published at KDD 2024 (ACM SIGKDD Conference on Knowledge Discovery and Data Mining), with the arXiv preprint 2311.09735 first posted in November 2023. The research team spanned Princeton University, IIT Delhi, and Georgia Tech — a cross-continental collaboration. In 2026, the paper continues to be cited on Princeton's research portal as foundational work in AI content optimization.

Core findings at a glance: Expert quotations boost AI visibility by +41%. Statistics addition delivers +33%. Fluency optimization adds +29%. Citing sources provides +28%. Keyword stuffing is the only tested strategy with negative impact at −8%. Pages ranked 5th in traditional search achieve up to +115% visibility lift after GEO optimization, while top-ranked pages lose up to 30%. The combined fluency + statistics strategy outperforms any single strategy by an additional +5.5%.

How GEO-bench measures AI visibility

GEO-bench is constructed from 10,000 real search queries drawn from 9 authentic datasets, including Google Search logs and Perplexity user queries. Each query is run through a generative search engine, and the research team analyzes which sources are cited in each AI answer and how prominently.

The core metric is position-adjusted word count: it counts how many words from a source appear in the AI-generated answer, with words appearing earlier in the answer weighted more heavily. This design distinguishes between "briefly mentioned" and "core reference" — a crucial distinction that simple binary citation metrics miss. The metric captures both whether you are cited and how prominently you are cited.

The 9 strategies: complete results table

The research team applied each content modification to a baseline page and compared visibility before and after on GEO-bench. Below are the complete results:

#StrategyVisibility liftBest for
1Expert quotations+41%Analysis, opinion, people
2Statistics addition+33%Law, policy, business
3Fluency optimization+29%Business, science, health
4Cite sources+28%Factual queries
5Quotation addition+21%History, biography
6Easy-to-understand language+11%Technical topics
7Technical jargon avoidance+9%General audience
8Authoritative tone+8%Trust-building
9Keyword stuffing−8%⚠️ Harmful in GEO

Source: Aggarwal et al., "GEO: Generative Engine Optimization," arXiv:2311.09735, KDD 2024. Visibility measured by position-adjusted word count on GEO-bench (10,000 queries × 9 datasets).

The most counterintuitive finding: lower rank, higher GEO returns

Perhaps the most strategically significant finding from the Princeton study is that GEO's impact is inversely correlated with traditional search rank. Pages ranked 5th in Google search achieved up to +115% visibility improvement after GEO optimization. Meanwhile, the 1st-ranked page experienced a 30% visibility decline.

This finding inverts the traditional SEO logic where top-ranking pages compound their advantage through authority signals. In AI search, the re-ranking models do not inherit Google's ranking conclusions — they evaluate content quality independently. BrightEdge's February 2026 analysis confirms this mechanism at scale: 62–83% of sources cited by Google AI Overviews come from outside the organic top 10. The implication is clear — pages that struggle in traditional SEO have the most to gain from GEO.

"For two decades, SEO was a winner-take-most game — high-authority domains with deep backlink profiles consumed the majority of traffic. GEO may break this pattern. If content quality is strong enough — original data, expert quotations, clear structure — a new domain can earn more exposure in AI answers than established authority sites."
— Interpretation of Princeton GEO findings, supported by BrightEdge Feb 2026 data

2026 independent validations of the Princeton findings

Since publication, multiple independent studies have validated and extended the Princeton team's findings. Here is a summary of the key validation studies as of September 3, 2026:

Study / SourceKey findingYear
Conductor AEO/GEO Benchmarks ReportFact density and citation sourcing show consistent positive correlation with AI visibility across hundreds of commercial sites2026
GrackerAI State of GEO 2026FAQPage Schema delivers 3.7× citation lift; 30-day content freshness earns 2.8× citation probability (BrightEdge/Semrush joint data)2026
BrightEdge AI Overviews Coverage Analysis62–83% of cited sources outside organic top 10; content published within 30 days has 2.8× citation probabilityFeb 2026
Authoritas GEO Signal StudyPages with original statistics see +156% citation probability in Google AI Overviews2026
Axis Intelligence ASFI™ IndexAI search market concentration dropped 50.5% in 12 months; multi-platform optimization increasingly critical2026
University of Toronto (arXiv:2509.08919)Brands cited on third-party platforms earn 6.5× more AI citations than self-hosted content alone2025
Semrush 2026 AI Visibility StudyAI search referral traffic grew 527% YoY; brands actively tracking GEO citations outpace peers as Google AI Overviews reach ~48% of US queries2026
Primores GEO/AEO Benchmarks 202696% of AI Overview citations come from sources with strong E-E-A-T signals; brand mentions correlate 3× more strongly with AI visibility than backlinks2026
Gartner Search Displacement Forecast25% of traditional desktop search volume will shift to AI chatbots and virtual agents by 2026 — validating GEO as a structural, not seasonal, discipline2024 (forecast)
GrackerAI State of GEO 2026GEO market sized at $7.3B, growing 34% CAGR; FAQPage Schema delivers 3.7× citation multiplier — corroborating the Princeton fluency + statistics lift at scale2026
Kumar et al. "GEO at Scale" (arXiv:2606.20065)First large-scale benchmark — 100K+ prompt responses across 100+ brands. Brand-stature ladder: household names cited in 73% of answers, mid-market 44%, niche brands 11%. 78% of citations go to corporate sites; ranked "best-of" listicles are the top format at 21%June 2026
Seer Interactive AI Overview CTR study2.43B impressions, 53 brands, 5.47M queries over 14 months: cited brands earn 120% more organic clicks per impression than uncited brands on the same AIO query (20,743 vs. 9,445 per million impressions) — the first large-sample monetary value of a citationApril 2026
Ahrefs citation position distributionOnly ~38% of AI Overview citations come from organic positions 1–10; 31% come from 11–100 and 31% from beyond position 100 — corroborating the Princeton "lower rank, higher return" finding independently2026
Promptwatch citation volatility trackingReddit's share of ChatGPT Search citations fell 86.4% in four days (3.83% → 0.52%) after a fan-out change, while Google AIO declined only 11.3% — proving citation share is controlled by the engine, not the publisherAugust 2026

The July 2026 critical survey: what 45 studies actually prove

A July 2026 critical survey (Martinez, arXiv:2607.14035) reviewed 45 GEO studies published between November 2023 and July 2026 — the first systematic audit of the field. Its uncomfortable but useful conclusion: GEO techniques reliably change how an already-retrieved page is cited, but no reviewed technique shows a stable, cross-platform causal effect on organic discoverability. In other words, the research shows how to be quoted better once an engine already has your page — not how to make the engine find you in the first place.

The survey also traces the famous "+40%" to its source. In Aggarwal et al. (KDD 2024), the metric is position-adjusted word count — the share of the generated answer attributed to a source, weighted by position. For the quotation tactic it rose from 19.3 to 27.2, about +41% in relative terms. That is a share-of-answer-text gain inside a fixed testbed, not a click gain and not a retrieval gain. The generalized claim that "GEO increases visibility by 40%" is explicitly listed among the claims the review rejects as unsupported.

What survives replication: topical relevance and early placement in the model context are the only levers with reproducible evidence. The survey reframes visibility as a stochastic, 7-stage pipeline — activation, crawling/indexing, retrieval, reranking, generation/citation, absorption/fidelity, and attention/conversion — where improving one stage says nothing about the others. Day-to-day source overlap is low (Jaccard 0.34–0.42), so a single before/after screenshot proves nothing without repeated sampling.

The practical takeaway is not that GEO is worthless — it is that indexing and retrieval are the real bottleneck, not formatting tricks. The Princeton benchmark remains the gold-standard method for measuring citation once a page is in the pool; the 2026 survey tells you not to confuse "more citable" with "more findable." This is exactly why the indexing workflow in our get-indexed-by-AI-search-engines guide matters as much as on-page optimization.

August 2026: the volatility evidence that reframes the +41%

Three weeks after that survey, the field got its sharpest natural experiment. Between July 18 and August 7, 2026, reddit.com held a steady 3.83% share of ChatGPT Search citations — one of the largest of any single domain. On August 14 it fell below 1%, and the August 14–17 average settled at 0.52%: an 86.4% relative drop in four days (Promptwatch). Two independent panels measured the same event: Qwairy recorded a 95% fall, and Peec AI recorded 88% — and found Reddit was not alone, with arXiv down 84% and YouTube down 78% over the same window.

The mechanism is what makes it a GEO lesson rather than a Reddit story. On August 8, ChatGPT's query fan-out changed: the share of fan-out queries using the site: operator jumped from 0.37% to 16.8% — roughly a 46× increase in a single day — while the average number of searches per response nearly doubled, from 1.08 to 1.83. The domain-scoped searches were added on top of the generic ones rather than replacing them. Crucially, Google's surfaces showed a different shape entirely: AI Overviews declined only 11.3% (2.37% → 2.10%) and AI Mode 30.5% (2.22% → 1.54%), both gradually, with no cliff. Reddit did nothing wrong, was being added to the S&P 500 that same week, and still lost almost its entire ChatGPT citation footprint in four days.

"Reddit's experience exposes GEO's original sin, the same one SEO has always had: you are building on ground you do not control. When a single model provider adjusts how it selects sources, your visibility can be rewritten overnight, with no notice and no explanation."
— Industry analysis of the August 2026 citation collapse, following Promptwatch, Qwairy, and Peec AI data

Read alongside the Princeton paper, this clarifies exactly what the +41% does and does not buy you. Princeton measures share of answer text, conditional on your page already being in the retrieved pool. The August collapse happened one stage earlier — at source selection, before any page-level optimization could matter. Both facts are true simultaneously, and confusing them is the most common strategic error in GEO today. The full breakdown is in our dedicated analysis, AI Citation Volatility 2026.

Practical application: a 5-step content checklist from the Princeton study

Based on the Princeton findings and their 2026 validations, here is a prioritized checklist for content optimization:

  1. 1.
    Add expert quotations (+41% lift) — Interview domain experts with full attribution. Use blockquote formatting with named sources. GrackerAI 2026 reports that attributed quotes increase AI citation probability by 3.7×, the single largest structured data multiplier.
  2. 2.
    Replace vague claims with statistics (+33% lift) — Every assertion should include a verifiable number with source and date. Authoritas 2026 found that pages with original statistical data achieve +156% higher citation probability in Google AI Overviews.
  3. 3.
    Optimize readability (+29% lift) — Remove redundancy, ensure clear subject-verb structure, limit sentences to 25 words. The fluency + statistics combination outperforms individual strategies by +5.5% (Princeton paper).
  4. 4.
    Cite authoritative sources (+28% lift) — Link every factual claim to academic papers, official statistics, or recognized industry reports. University of Toronto 2025 research (arXiv:2509.08919) found that brands cited on third-party platforms receive 6.5× more AI citations than those appearing only on owned domains.
  5. 5.
    Stop keyword stuffing (−8% penalty) — AI engines use semantic embedding matching, not keyword counting. BrightEdge 2026 data confirms natural writing outperforms keyword-optimized content for AI citation. Use synonyms and related entities instead.

September 2026: the failure-mode taxonomy that reorders the nine tactics

The most useful thing published about the Princeton tactics in 2026 is not a replication — it is a study asking a different question. Instead of "which edit produces the biggest lift in a controlled test," the AgentGEO paper asked: when content fails to get cited at all, what is actually going wrong? The answer reframes the whole nine-strategy list.

Failure categoryShare of citation failuresWhat it covers
Semantic alignment62.2%Intent divergence, contextual gap, outdated information, localization mismatch
Content quality27.1%Information scarcity, unstructured layout — the bucket the Princeton tactics target
Technical integrity10.1%Access blocking, JS rendering, unparseable content, low signal-to-noise
Systemic exclusion0.6%Competitive redundancy, context-window truncation

Source: AgentGEO academic taxonomy of AI citation failure modes (2026). Read alongside the Princeton nine-strategy list, which is a content-quality intervention set.

Put the two side by side and the mismatch is obvious. Statistics addition and source citation are content-quality interventions, so they can only ever operate inside the 27.1% slice — and even there they are one sub-tactic, not the category. The single largest reason content is skipped is semantic alignment: the page answers a question nobody is currently asking, in the wrong framing, or with data that has gone stale. No number bolted onto a paragraph fixes that.

The same 2026 literature also clarifies where Princeton's testbed sits in the pipeline. Sequenced, four decisions happen before any page-level tactic can matter:

  1. 1.
    Does the engine search at all? Profound found Claude invokes live web search for only 36.6% of tested prompts. For the rest there is no citation opportunity to win, regardless of how well the page is optimized.
  2. 2.
    Which sources get selected from what is retrieved? DEJAN measured OpenAI selecting Reddit as a source in just 0.61% of retrieved candidates after the August 2026 change — a selection decision, not a content-quality one.
  3. 3.
    Why does retrieved content fail to get cited? This is AgentGEO's stage, and it is dominated by semantic alignment at 62.2%.
  4. 4.
    Was the recommendation already set? Seer Interactive's ghost-citation work found a roughly 5× citation-rate lift once a brand is already the chosen recommendation — meaning some citation decisions were effectively settled before retrieval began.
"Princeton tells you what a page-level edit is worth once the page has cleared every earlier stage. AgentGEO tells you that most pages never clear those stages. Both are correct, and reading only the first one is why so much GEO effort lands on the wrong 10% of the problem."
— GeoAura Research, synthesis of the Princeton KDD 2024 benchmark and the 2026 AgentGEO failure taxonomy

None of this makes the nine tactics wrong — it makes them conditional. Applied to a page that is retrieved, semantically aligned, and in the consideration set, statistics (+33%), quotations (+41%), and citations (+28%) remain the best-measured levers available. Applied to a page that fails stage 1 or stage 3, they change nothing. Diagnose the stage first, then spend the tactic.

Two further 2026 data points belong in that diagnosis. A ConvertMate benchmark across 12,500 queries and 8,000 domains found 83% of AI Overview citations go to pages outside the organic top 10, with AI search traffic converting at roughly 4.4× organic — the upside for pages that do clear retrieval is real. And Ahrefs' 75,000-brand analysis found unlinked brand mentions correlating 0.656–0.709 with AI visibility while raw backlinks sat near zero, which is the off-site counterpart to the same lesson: being retrieved and being named are upstream of being well-formatted.

Study limitations

While the Princeton study is the methodological gold standard in GEO, several limitations merit attention:

  • The experiment used a single generative engine (an open-source GPT-3.5 system), not production ChatGPT Search or Perplexity. Different AI engines may exhibit different citation behaviors.
  • The 9 strategies were tested individually, not in combination (except the fluency + statistics pair). In practice, multiple strategies interact, and the combined effects may differ from individual lifts.
  • 10,000 queries, while substantial, may not represent highly niche or industry-specific search intents. Results may vary by vertical.
  • The benchmark predates the ChatGPT Search launch (October 2024), Google AI Mode (March 2025), and the 2026 shifts in market dynamics. Subsequent validations address these gaps but do not replicate the controlled experimental design.

The directional conclusions should be treated as reliable guidance rather than precise prediction formulas. The convergence of independent validations — Conductor, GrackerAI, BrightEdge, Authoritas, Axis Intelligence — provides confidence in the overall framework while acknowledging that exact lift percentages vary by platform, vertical, and content type.

Frequently asked questions

What is the Princeton GEO study?

It is the paper "GEO: Generative Engine Optimization" by Aggarwal, Dugan, et al., published at KDD 2024 (arXiv:2311.09735). The core contribution is GEO-bench — 10,000 queries across 9 datasets — used to systematically test 9 content optimization strategies for AI search visibility. As of September 2026, it remains the most cited foundational benchmark in the GEO field, and the only GEO term with peer-reviewed grounding.

What is GEO-bench and how does it measure visibility?

GEO-bench contains 10,000 authentic search queries drawn from 9 datasets, including Google Search logs and Perplexity user queries. Visibility is measured via position-adjusted word count — the number of words from a source appearing in an AI answer, weighted by position — plus a subjective impression score. Both metrics are computed on documents already placed in the model's context.

Which GEO strategy gives the biggest visibility lift?

Expert quotations (+41%) lead all nine strategies. Statistics addition (+33%), fluency optimization (+29%), and citing sources (+28%) complete the top tier. Authoritative tone adds only +9%, and keyword stuffing is uniquely harmful at −8%. GrackerAI 2026 data further shows FAQPage Schema produces a 3.7× citation multiplier.

Does GEO work better for low-ranking or high-ranking pages?

Low-ranking pages gain far more. Pages ranked 5th in traditional search achieved up to 115% visibility improvement after GEO optimization, while top-ranked pages lost 30%. Independent 2026 data points the same way: BrightEdge reports 83% of Google AI Overview citations come from outside the top 10, and Ahrefs found only ~38% of AI Overview citations come from organic positions 1–10.

Are the Princeton GEO study findings still valid in 2026?

Directionally yes, with one important correction. Conductor, GrackerAI, BrightEdge, and Authoritas all validate the core conclusion that fact density, citation sourcing, and expert authority drive AI citation. But the July 2026 critical survey of 45 studies found the popular "GEO increases visibility by 40%" generalization is unsupported: the +40% is a share-of-answer-text gain inside a fixed testbed, not a click or retrieval gain.

What is the Princeton GEO study citation on collaborate.princeton.edu?

Princeton maintains a publication page for the paper on collaborate.princeton.edu, the university's research collaboration portal, alongside the arXiv preprint at arXiv:2311.09735. For academic citation, use the KDD 2024 conference version: Aggarwal, P., Dugan, L., et al., "GEO: Generative Engine Optimization," Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024. The arXiv preprint is the most commonly linked open-access version.

Does the +41% quotation lift translate into more clicks?

Not directly. Princeton's metric is position-adjusted word count — share of answer text, which for the quotation tactic rose from 19.3 to 27.2 (+41% relative). That is a share-of-answer gain inside a fixed testbed, not a click or retrieval gain. Citation does carry measurable click value though: Seer Interactive (April 2026) found cited brands earned 120% more organic clicks per impression than uncited brands on the same AI Overview query, and 93% of Google AI Mode sessions still end without any click at all.

What happened to Reddit's ChatGPT citations in August 2026?

Reddit held a steady 3.83% share of ChatGPT Search citations from July 18 to August 7, then fell below 1% on August 14, averaging 0.52% through August 17 — an 86.4% relative drop in four days (Promptwatch). Qwairy measured 95% and Peec AI 88%. The trigger appears to be a fan-out change on August 8 in which site:-scoped searches jumped from 0.37% to 16.8%. Google AI Overviews declined only 11.3% in the same window.

What share of citation failures do the Princeton tactics actually address?

About a quarter. The AgentGEO taxonomy, the first systematic classification of why AI systems skip content, attributes 62.2% of citation failures to semantic alignment (intent divergence, contextual gap, stale information, localization mismatch), 27.1% to content quality, 10.1% to technical integrity, and 0.6% to systemic exclusion. Statistics and citations are content-quality interventions — they live entirely in the 27.1% slice. No bolted-on statistic fixes a page answering a question nobody is asking.

Do the Princeton tactics help if an engine never retrieves the page?

No, and this is the most common strategic misread of the study. Every tactic is measured on documents already placed in the model's context, so none addresses whether a page survives retrieval and reranking. That upstream stage is where most losses occur: Profound found Claude invokes live web search for only 36.6% of tested prompts, and a 2026 ConvertMate benchmark across 12,500 queries and 8,000 domains found 83% of AI Overview citations go to pages outside the organic top 10, with AI search traffic converting at roughly 4.4× organic.

References: Aggarwal, P., Dugan, L., et al. "GEO: Generative Engine Optimization." arXiv:2311.09735, KDD 2024. · Princeton University, IIT Delhi, Georgia Tech. · Conductor 2026 AEO/GEO Benchmarks Report. · GrackerAI "State of Generative Engine Optimization: 2026 Data Sheet" (Gartner, McKinsey, Forrester, Princeton, eMarketer, 30+ sources). · BrightEdge Feb 2026 AI Overviews Coverage Analysis. · Authoritas 2026 GEO Signal Study — original statistics +156% citation probability. · Axis Intelligence "AI Search Statistics 2026" — ASFI™ market concentration index. · Primores "GEO/AEO Benchmarks 2026" — 96% of AIO citations from strong E-E-A-T sources, brand mentions 3× backlinks for AI visibility. · University of Toronto "Third-Party Citation Multiplier" arXiv:2509.08919 (2025). · Digital Applied "AI Search Engine Statistics 2026: Market Share Data." · collaborate.princeton.edu — GEO publication page (accessed September 3, 2026).

Want to check your site's GEO readiness?

Run the 27-point GEO audit

Related articles

What Is GEO (Generative Engine Optimization)? Complete Guide

Updated September 2026: GEO is the practice of optimizing content to be cited and referenced by AI search engines like ChatGPT Search, Perplexity, Google AI Overviews, Gemini, and Claude. New this month: Ahrefs Brand Radar counted 462.8M prompts per month across six AI surfaces (AI Overviews 308.3M, Gemini 31.5M, Perplexity 31.4M, ChatGPT 31.3M, Copilot 30.9M, AI Mode 29.4M), SparkToro measured 68.01% of Google searches ending without a click, a 95-domain probe found a median text-to-HTML ratio of just 2.4%, and a September 7 2026 replication found volume-controlled quotation, statistics and citation edits produced no pooled lift - while adding query relevance tripled the predictive score from 0.114 to roughly 0.37. Plus the Princeton KDD 2024 strategy lifts, the 75,000-brand mention-versus-backlink inversion, and a full multi-platform optimization framework.

GEO vs SEO: 7 Critical Differences You Need to Know (2026 Update)

SEO targets keyword rankings and clicks. GEO targets AI citations and brand mentions. With AI search traffic growing 527% YoY in 2026, Google AIO covering 48-50% of queries with 62-83% of sources outside organic top 10, and Gartner predicting 25% search volume decline, this guide breaks down the 7 key differences with fresh 2026 data and verified statistics — and as of late August 2026, 32% of marketing leaders rank GEO their #1 2026 priority (BrightEdge).

How AI Search Engines Work: RAG Architecture Explained

Updated September 2026: Google confirmed there is no separate AI index for AI Overviews and AI Mode and told publishers to deprioritize AEO/GEO hacks such as content chunking. Three new studies qualify the classic four-stage RAG model: a replication found no positive pooled effect for quotations (-0.325 pp), statistics (-0.276 pp) or source citations (-0.793 pp); Trellner found 59.8% of Perplexity citations come from domains ranked worse than #100,000; and Intender found only 8% source overlap across five engines. Complete 4-stage pipeline breakdown with stage-by-stage optimization for ChatGPT, Perplexity, Gemini, Claude and Google AI Mode.