Entity Density for AI Citations: The 20.6% Benchmark and the Test That Broke It
Entity density is the most-quoted number in GEO that has never survived a controlled test. Kevin Indig's analysis of 1.2 million ChatGPT responses puts heavily cited content at 20.6% entity density against 5-8% for standard English, and pages with 15+ recognized entities show 4.8x higher Google AI Overview selection probability. Then the GEO Lab ran the only controlled test — 10 pages spanning 12.9 to 29.5 unique entities per 1,000 words across 22 pre-registered Perplexity queries — and got zero citations at every level, because the domain never cleared entity recognition. Covers why recognition gates density, Google's 0-1 salience scale, the 3-billion-entity Knowledge Graph cleanup of 2025, GPT-6 Astra's agentic browsing, and a 6-step method for raising density honestly.
Entity density is the most-quoted number in GEO that has never survived a controlled test. Kevin Indig's analysis of 1.2 million ChatGPT responses puts heavily cited content at 20.6% entity density against 5-8% for ordinary English text. Then the only experiment that actually varied density -- the GEO Lab's April 2026 test on Perplexity, 10 pages spanning 12.9 to 29.5 unique entities per 1,000 words, 22 pre-registered queries -- returned zero citations at every single level. Both findings are true. The reconciliation is the useful part: density multiplies recognition, it does not create it.
The evidence in one place (2026): · 20.6% entity density in heavily cited content vs 5-8% in standard English (Kevin Indig, 1.2M ChatGPT responses). · 4.8x higher Google AI Overview selection probability for pages carrying 15+ recognized entities. · 36.2% vs 20.2% — citation-winning content is roughly twice as likely to use definitive language such as "is defined as". · Zero citations across all 22 queries in the only controlled density test (GEO Lab, April 2026, Zenodo DOI 10.5281/zenodo.19450361). · 96% of AI Overview citations come from sources that clear E-E-A-T credibility thresholds (ZipTie.dev).
What entity density actually measures
An entity is a specific, identifiable thing with attributes and relationships. A keyword is a string. That distinction is the whole basis of the practice: Google drew the line publicly on May 16, 2012, when it launched the Knowledge Graph with 500 million objects and 3.5 billion facts under the slogan "things, not strings." Practitioners now estimate the graph holds more than 5 billion entities and 500 billion facts, though Google publishes no current count1.
Entity density is simply the share of your words that are named things rather than filler. "Our platform helps teams work better" contains zero entities. "GeoAura tracks citations across ChatGPT, Perplexity, Google AI Overviews, Gemini and Claude" contains five. Measured mechanically, that is the gap between a page an engine can place in a knowledge neighbourhood and one it cannot.
| Text type | Entity density | What it signals |
|---|---|---|
| Standard English prose | 5-8% | Baseline. Nothing for a retrieval system to anchor on. |
| Heavily AI-cited content | 20.6% | Observational benchmark from 1.2M ChatGPT responses2. |
| Controlled test range | 12.9-29.5 / 1,000 words | The span that produced zero citations3. |
| Pages with 15+ entities | 4.8x selection | Higher probability of Google AI Overview inclusion2. |
The controlled test that returned nothing
This is the part almost every entity-density guide leaves out. In April 2026 the GEO Lab ran the cleanest test of the hypothesis available: 10 pages from a single domain, deliberately spanning 12.9 to 29.5 unique entities per 1,000 words, measured with a filtered spaCy named-entity pass that was locked before results were examined. Artur Ferreira pre-registered 22 competition queries where three to five of those pages counted as plausible answers.
Every query, every page, every density level returned zero citations. The correlation came back undefined for lack of variance. The preprint sits on Zenodo under DOI 10.5281/zenodo.19450361 with raw data on GitHub3.
"The same domain hit a 20% citation rate on a separate 30-query set — 6 citations, all in first position. Every one landed on a concept the site originated."
Read that second result carefully, because it is the actual finding. The domain did get cited — 20% of the time, all at position one — but only on concepts it originated. What earned those citations was not density. It was being the only source that could corroborate a specific claim. The same limitation applies to the foundational study everyone cites: Aggarwal et al. (KDD 2024) measured up to 41.5% visibility gains from adding citations, but their manipulations ran on pages already inside the retrieval candidate set. That bounds the result to domains that have already cleared recognition.
Why recognition gates density
Retrieval systems run an identity pass before a relevance pass. They match the query to entities, pull candidate passages, then check whether independent sources corroborate the claim before generating an answer. Machine Relations synthesized more than 680 million AI citations in May 2026 and described the mechanism plainly: AI engines select sources rather than rank pages, and the selection test is whether a brand appears across multiple independent contexts the engine can cross-reference4.
Peer-reviewed retrieval research supports the mechanic. Work on uncertainty-driven evidence selection shows retrieval-augmented systems favour evidence that reduces uncertainty, not evidence that maximizes relevance alone. An organization named in one source carries high uncertainty; the same organization named across five independent sources carries low uncertainty and wins the slot. ZipTie.dev's analysis found 96% of AI Overview citations come from sources clearing E-E-A-T credibility thresholds5. A fragmented brand identity raises assembly cost, and the engine cites the competitor it can verify faster.
The 2025 Knowledge Graph cleanup raised the bar
- June 2025: Google removed more than 3 billion entities in two closely timed updates — roughly twice the net additions of the entire previous year. The "event" category dropped 76.91%.
- August 2025: a second cleanup focused on corporation, organization and brand entities.
- March 2026 core update: continued the trajectory, strengthening E-E-A-T and entity-based authority.
- Implication: Google traded volume for confidence. Mentions alone no longer register; entities must be well defined and consistently described6.
Salience: why naming more entities is not the goal
The Google Natural Language API scores entity salience on a 0 to 1 scale — how central each entity is to the page's main topic. This is where most entity-density advice goes wrong. Stuffing thirty brand names into a paragraph raises density and lowers salience on all of them, because the model cannot tell what the page is about. Google's Enterprise Knowledge Graph formally supports only three entity types for linking — Organization, LocalBusiness and Person — so make those three unambiguous before chasing anything exotic7.
There is a measurable stylistic correlate. Citation-winning content is almost twice as likely (36.2% vs 20.2%) to contain definitive language such as "is defined as" or "refers to", which naturally introduces a named entity with clear attributes. Definition sentences do two jobs at once: they raise salience on the primary entity and they hand the model a liftable triple2.
How to raise entity density honestly: 6 steps
- 1.Clear the recognition gate first. Before touching density, confirm the engine can resolve you at all: a Knowledge Panel, a Wikidata entry, consistent naming across LinkedIn and Crunchbase. If the identity pass fails, no amount of density matters. See how to get cited by AI search.
- 2.Open with a definition sentence. "Entity density is the share of words that are named entities" beats any build-up. It raises salience on your primary entity and gives the model a liftable fact in the first 30% of the page, where 44.2% of LLM citations originate.
- 3.Name the specific things, not the category. Write "OAI-SearchBot" not "the AI crawler"; "Gemini 3.8 Flash, deployed 2 September 2026" not "the latest model". Specificity is what turns a mention into an edge.
- 4.State relationships as sentences. GraphRAG-style retrieval builds an entity graph from the corpus first. "Conductor analysed 167,867,680 AI Overview citations" contributes a real edge; an implied connection contributes an isolated node.
- 5.Mirror it in sameAs, and let the prose stand alone. Point Organization and Person schema at Wikidata, LinkedIn and Crunchbase, but never rely on markup to carry a fact the prose does not state — models ground on rendered text, not on JSON-LD they never receive.
- 6.Originate something. The GEO Lab's 20% citation rate landed entirely on concepts the site originated. Publishing a measurement nobody else has is the only entity signal that reliably survives the corroboration test.
How to measure your own entity density
| Method | How | Best for |
|---|---|---|
| spaCy NER pass | Run a filtered named-entity recognition pass over rendered text; count unique entities per 1,000 words | Reproducible internal benchmarking — lock the config before you look at results |
| Google NL API | Returns entity list with a 0-1 salience score each | Checking whether your intended primary entity is actually dominant |
| Entity coverage audit | List every entity a topic expects and check which your page names | Finding gaps density scoring cannot see |
| Cross-engine recall test | Ask each engine "what do you know about [brand]" and "who are the leading experts on [topic]" | The recognition gate itself — the test that actually decides outcomes |
"The web is a graph of entities and their relationships. Pages are just the surface through which those relationships are expressed."
What changed in September 2026
OpenAI released GPT-6 Astra on September 3, 2026 with multi-step web research built into the default reasoning path rather than exposed as an opt-in tool. Where earlier ChatGPT models relied on cached index snippets, Astra opens pages and reads them. That makes clean named-entity definitions a direct retrieval asset: a page that states what it is gives the model something to lift, and a page that relies on context fails the test8.
Two boundaries matter. First, 87% of ChatGPT Search citations match Bing's top-10 organic results, so classic index position still predicts eligibility — entity work raises your odds within the candidate pool, it does not create one. Second, ChatGPT mentions brands 3.2 times more often than it cites them with an attributed link, so entity clarity is doing more work for unlinked mentions than for clicks.
Frequently asked questions
What is entity density in content optimization?
Entity density is the share of words on a page that are named entities — specific people, organizations, products, places, standards, crawler names or dated versions — rather than generic nouns and filler. It is measured with a named-entity recognition pass over the rendered text, most commonly spaCy or the Google Natural Language API, and reported either as a percentage of all words or as unique entities per 1,000 words.
What entity density should I target for AI citations?
Kevin Indig's analysis of 1.2 million ChatGPT responses found heavily cited content sits at roughly 20.6% entity density against 5-8% for standard English text. Treat 20% as an observational benchmark, not a target to hit by force: the only controlled test of entity density found zero citations at every level from 12.9 to 29.5 unique entities per 1,000 words, because the test domain had not cleared entity recognition.
Does higher entity density guarantee more AI citations?
No. The GEO Lab ran a controlled experiment on Perplexity in April 2026 across 10 pages spanning 12.9 to 29.5 unique entities per 1,000 words, with 22 pre-registered competition queries. Every query, page and density level returned zero citations, and the correlation came back undefined for lack of variance. Density is a multiplier on recognition, not a substitute for it.
Why did the controlled entity density test return zero citations?
Because the domain never cleared entity recognition in the first place. Retrieval systems run an identity pass before a relevance pass: they match the query to known entities, pull candidate passages, then check whether independent sources corroborate the claim. Machine Relations synthesized more than 680 million AI citations and describes the test as cross-context corroboration, not ranking. A domain no source corroborates is never a candidate regardless of how dense its prose is.
How do AI engines actually use entities?
In two passes. During training, a model that repeatedly sees a brand near a stable set of concepts learns a tight association between them, so entities with thin corpus presence get retrieved less often regardless of quality. At inference time, retrieval-augmented systems fetch live pages and extract entities from them, and the pages selected for citation are usually the ones whose entity coverage matches the query intent rather than the ones with the highest classic ranking score.
What is entity salience and why does it matter?
Salience is how central an entity is to the page's main topic, scored by the Google Natural Language API on a 0 to 1 scale. It matters because naming many entities is not the same as being about one. A page that mentions thirty brands in passing has low salience on all of them; a page whose primary entity appears in the title, opening definition and every section scores high and is far easier for an engine to attribute.
Does schema markup help entity recognition?
It helps disambiguate, but it does not substitute for prose. Models ground on rendered text they can read, not on JSON-LD they never receive, so the sentence has to carry the fact on its own. Use sameAs arrays to point Organization and Person entities at Wikidata, LinkedIn and Crunchbase profiles, and keep the prose unambiguous about which entity it means.
Which entity types does Google formally recognize?
Google's Enterprise Knowledge Graph currently supports three primary entity types for formal linking: Organization, LocalBusiness and Person. That is a narrow set, so prioritize getting those three unambiguous before investing in exotic types. The graph itself launched on May 16, 2012 with 500 million objects and 3.5 billion facts under the phrase things, not strings, and practitioners now estimate it holds more than 5 billion entities.
How did the 2025 Knowledge Graph cleanup change entity SEO?
It raised the bar. In June 2025 Google removed more than 3 billion entities in two closely timed updates, roughly twice the net additions of the entire previous year, with the event category dropping 76.91%. A second cleanup in August 2025 focused on corporation, organization and brand entities. Google traded volume for confidence, so merely mentioning an entity is no longer enough — it has to be well defined and consistently described.
Does GPT-6 Astra change what entity signals matter?
It sharpens them. OpenAI released GPT-6 Astra on September 3, 2026 with multi-step web research built into the default reasoning path, which means the model opens your page and reads it rather than relying on cached snippets. A page that opens with a clean named-entity definition gives it something to lift; a page that relies on context to establish what it is about fails that test. Note also that 87% of ChatGPT Search citations match Bing's top-10 organic results, so index position still predicts eligibility.
References:
1 Digital Applied (May 2026), Knowledge Graph size estimates; Google Knowledge Graph launch, 16 May 2012 — 500 million objects, 3.5 billion facts, "things, not strings".
2 Kevin Indig, analysis of 1.2 million ChatGPT responses (2026) — heavily cited content at 20.6% entity density vs 5-8% in standard English; pages with 15+ recognized entities show 4.8x higher AI Overview selection probability; citation-winning content 36.2% vs 20.2% likely to use definitive language. Via RankDraft.
3 GEO Lab / Artur Ferreira, controlled entity-density experiment on Perplexity (April 2026) — 10 pages, 12.9 to 29.5 unique entities per 1,000 words, filtered spaCy NER locked before analysis, 22 pre-registered competition queries; zero citations at every level. Preprint: Zenodo DOI 10.5281/zenodo.19450361.
4 Machine Relations (May 2026), synthesis of 680M+ AI citations — engines select sources rather than rank pages; selection test is cross-context corroboration.
5 ZipTie.dev, AI Overview citation analysis — 96% of citations come from sources clearing E-E-A-T credibility thresholds.
6 Knowledge Graph cleanup reporting (June and August 2025) — 3B+ entities removed; "event" category down 76.91%; second pass on corporation, organization and brand entities.
7 Google Cloud Natural Language API — entity salience scored 0 to 1; Google Enterprise Knowledge Graph formally supports Organization, LocalBusiness and Person for entity linking.
8 OpenAI, GPT-6 Astra release notes (3 September 2026) — multi-step web research in the default reasoning path; 87% of ChatGPT Search citations match Bing top-10 organic results; ChatGPT mentions brands 3.2x more often than it cites them.
9 Aggarwal et al., "GEO: Generative Engine Optimization," arXiv:2311.09735, KDD 2024 — up to 41.5% visibility gains from citations, measured on pages already inside the retrieval candidate set.
Want to check your site's GEO readiness?
Run the 27-point GEO auditRelated articles
9 Proven GEO Optimization Strategies (With Quantified Data)
Expert quotations boost AI visibility by 41%, statistics by 33%, fluency by 29%, citations by 28% — and statistics + citations compound to ~+61%. Updated August 2026 with validation from Conductor, GrackerAI, BrightEdge, Authoritas, and Previsible, plus the GEO market now worth $7.3B at 34% CAGR. Updated Aug 2026: ChatGPT passed 1B MAU and 32% of marketing leaders rank GEO their top 2026 priority (BrightEdge). The complete peer-reviewed guide to all 9 GEO strategies (Princeton, KDD 2024) with quantified lift percentages.
How to Get Cited by AI Search: The 2026 Playbook
Getting cited by ChatGPT Search, Perplexity, Google AI Overviews, Gemini, and Claude is a repeatable process, not luck. Updated September 2026: Ahrefs analysed 3 million+ US queries and found YouTube takes 22.9% of AI Overview citations (Reddit 18.5%, Facebook 10.1%), while LLM Pulse found 64.7% of citations on non-branded shopping queries go to brand and manufacturer sites. Also new this month: Google AI Mode entity-wrapped citations appear on 87.6% of recommendation prompts but 0% of informational ones, and a September 7 2026 replication found volume-controlled quotation and statistics edits produced no pooled lift. This 2026 playbook covers the 7-step citation workflow, the off-site reality that earned media drives 84% of AI citations (Muck Rack, May 2026), the 3.5B+ weekly AI queries you compete for, and the structural fixes (AI crawler access, Schema.org, FAQ) that move pages into AI answers.
AI Search Optimization: The Complete 2026 Guide
Updated August 2026: AI search optimization is the practice of structuring content so ChatGPT Search, Perplexity, Google AI Overviews, Gemini, and Claude cite it. AI search traffic grew 16× from 2024 to 2026, yet only ~12% of AI citations match Google top 10. This 2026 playbook covers the 9 GEO strategies with measured lift (+41% quotations, +33% statistics, +28% citations, +29% fluency), the 7-step optimization process, common mistakes (keyword stuffing −8%), and how to measure AI visibility.