Black-Hat GEO & AI Corpus Poisoning: The First Standards and What Gets You Delisted (2026)
GEO stopped being unregulated in 2026. Updated September 22 2026: T/CAPT 026-2026 (published 11 August 2026) names four prohibited practices — corpus poisoning, answer hegemony, pseudo-consensus manufacturing and prompt injection — and grades sources A/B/C/D, with A-tier content cited 7.2x more than C-tier commercial pages and more than 23x more than D-tier. Five group standards published 14 September 2026 by 41 organizations and 75 experts define seven pricing units, five pricing models and nine core metrics with minimum sample sizes, ban absolute-effect promises, and require 100% human review in finance and medical. Also new: Counter-GEO-Bench (arXiv, 2 September 2026, EMNLP 2026) shows a brand can be cited beside misinformation it never published, with tested safety filters cutting attack success by only 5.7%; the April 2026 CAC campaign classified GEO manipulation as AI data poisoning; a five-department rule took effect 1 September 2026 with 11 red lines; and CNNIC counted 602M generative-AI users at 42.8% penetration against 200+ vendors of which only about 19% have in-house technology. Includes the where-legitimate-GEO-ends table and a 6-point vendor audit.
GEO spent two years as an unregulated market, and 2026 is the year that ended. Between April and September 2026 a regulator classified GEO manipulation as a form of AI data poisoning, a self-discipline standard named four prohibited practices, a five-department rule turned content distribution into a legal duty, and five group standards defined how GEO may be priced, measured and audited. If you buy GEO services or publish at scale, the question is no longer only whether a tactic works — it is whether it survives an audit.
The reason regulators moved is that the tactics work often enough to matter. Counter-GEO-Bench (arXiv, 2 September 2026, accepted to EMNLP 2026) demonstrated that a brand can be cited next to misinformation it never published, and that the safety filters it tested cut attack success by only 5.7%. At the same time, the same benchmark shows why the honest tactics still win on the merits: the Princeton GEO study (KDD 2024) measured +33% from statistics, +41% from expert quotations and +28% from authoritative citations, while keyword stuffing costs about 8%. The gap between those two lists is the whole subject of this guide.
The 2026 rulebook in seven numbers: 4 prohibited practices named in T/CAPT 026-2026 (11 Aug 2026) · 11 red lines in the five-department distribution rule (1 Sep 2026) · 5 group standards, 41 drafting organizations and 75 experts (14 Sep 2026) · 7 pricing units and 9 core metrics with minimum sample sizes · A-tier sources cited 7.2x more than C-tier and 23x more than D-tier · 602M generative-AI users at 42.8% penetration (CNNIC) · 200+ GEO vendors, ~19% with real in-house technology.
What corpus poisoning actually is
Corpus poisoning is not a new version of link spam. The target has moved. Classic spam wanted a position in a list of ten blue links; poisoning wants to change what the model says. Because generative engines retrieve live pages and then synthesise, the attacker does not need to rank — they need to be one of the few corroborating sources a retrieval pass finds, and to agree with enough other planted sources to look like consensus.
In practice it takes four shapes, and the Chinese standards name all four:
- 1.Corpus poisoning (data injection)
Batch-injecting invented product parameters, fabricated partnership cases and forged certifications into low-authority sites so the model absorbs them as fact. Vendors marketing this as "brainwashing the AI" or "top-of-category recommendation without a product" are describing this practice. The Cyberspace Administration of China has formally characterised it as AI data poisoning.
- 2.Answer hegemony
Occupying every slot in the answer so no competing source can enter the retrieved set — typically by publishing the same claim across many domains under different bylines.
- 3.Pseudo-consensus manufacturing
Creating the appearance of independent agreement: dozens of near-identical "reviews", "tests" or "expert roundups" that a retrieval system reads as corroboration. This is the practice most likely to fool a cross-verification pass, and the one platforms are now actively discounting.
- 4.Prompt injection attacks
Embedding instructions in page content aimed at the agent reading it, rather than at the human — hidden text, vector-level manipulation, or directives designed to override the retrieval step.
A fifth pattern sits just outside the named four but is treated the same in practice: volume flooding — "unlimited keywords for a year" packages that cover long-tail questions with template-spun pages. That is not poisoning in the strict sense, but it produces the same failure mode, and platforms have started penalising it directly. DeepSeek has systematically reduced citation weight for recently published, highly homogeneous marketing listicles, and Kimi has rolled citation sources back to pre-2025 material to avoid manipulation entirely.
The 2026 enforcement timeline
Four instruments landed in five months. Read them as a sequence: a criminal-style classification, then a technical standard, then a legal duty, then a commercial standard.
| Date | Instrument | What it established |
|---|---|---|
| March 2026 | GEO industry self-discipline convention (10+ founding companies) | First collective commitment; banned one-click distribution of fake reviews, fake science explainers and fake authoritative reports |
| April 2026 | CAC "Qinglang" campaign on AI application abuses | GEO malicious marketing classified as AI data poisoning; enforcement rather than guidance |
| 11 Aug 2026 | T/CAPT 026-2026, first GEO self-discipline standard | Four prohibited practices; A/B/C/D source-tier grading; white-hat versus black-hat boundary drawn for the first time |
| 1 Sep 2026 | Five-department multi-channel content distribution rule | Distribution duties rise from self-discipline to legal obligation, with 11 red lines |
| 14 Sep 2026 | Five GEO group standards (China Advertising Association) | Terminology, vendor evaluation, pricing and measurement, trusted corpus, compliance — 7 pricing units, 5 pricing models, 9 core metrics |
"These standards end the state of GEO services having no standard to stand on. They give the market a common technical language and a basic floor of consensus — on terminology, on pricing, on what counts as a measurable result, and on which practices are out of bounds."
The source-tier system: why five A-tier pages beat fifty C-tier pages
The most useful thing T/CAPT 026-2026 produced is not the list of bans. It is a grading system that turns "quality" into something you can audit, and it comes with monitoring data attached.
| Tier | What counts | Relative AI citation rate |
|---|---|---|
| A | Government bodies, national authoritative databases, original state-media reporting | Highest — 7.2x C-tier |
| B | Academic papers, industry white papers, original mainstream media reporting | High |
| C | Corporate websites, commercial reports — supporting reference only | Baseline (1x) |
| D | Anonymous sources, low-quality aggregators — AI systems told not to rely on them | More than 23x below A-tier |
This is the same conclusion the Western data reached by a different route. Our analysis of Conductor 167,867,680 AI Overview citations found the position-one result is cited only 24.9% of the time and 57% of citations land outside the organic top 10 — ranking is not the gate. Ahrefs put the commercial version of the same point at 75,000 brands: mentions and backlinks invert as predictors. What the tier system adds is a procurement instruction: do not ask how many pages a vendor will publish, ask which tier they land in.
The uncomfortable part: poisoning works, and the penalties are asymmetric
Anyone selling GEO services should be honest about why the black-hat market exists. Three findings explain it.
- ▸ Injection succeeds more often than it should. Counter-GEO-Bench (arXiv, 2 Sep 2026) showed a brand being cited alongside claims it never made, with tested safety filters reducing attack success by only 5.7% — a defensive gap, not a solved problem.
- ▸ Retrieval is the soft target, not wording. Trellner found 59.8% of Perplexity citations come from domains ranked worse than #100,000, and our own cross-engine work puts URL overlap at 10.2% across engines. A source nobody ranks can still be retrieved — which is precisely the opening poisoning exploits.
- ▸ The penalty outlives the campaign. Once a cross-verification pass exposes a fabricated endorsement, the damage is not confined to a page: the entity record — name, address, organisation — is marked unreliable, and recovery costs far exceed what the campaign cost.
And the liability does not transfer. Under the Internet Advertising Measures, the advertiser is responsible for the truthfulness of the content, and that responsibility is not discharged by delegating the work to a third party. "The agency did it" is not a defence. Regulators have also signalled the reverse test: more than 70% of agencies surveyed were found to have a blurred understanding of where GEO boundaries sit, routinely repackaging bulk low-quality distribution as compliant GEO.
Where legitimate GEO ends and black-hat GEO begins
The boundary is drawn on verifiability and intent, not on technique. Nearly every legitimate GEO tactic has a black-hat twin that looks identical from the outside, and telling them apart is now a procurement skill.
| Legitimate tactic | Measured effect | The banned twin |
|---|---|---|
| Publishing original statistics | +33% AI visibility (Princeton KDD 2024); +156% AIO citation probability (Authoritas) | Inventing the statistic and injecting it into low-authority sites |
| Adding expert quotations | +41% AI visibility (Princeton KDD 2024) | Fabricating the expert, the institution or the endorsement |
| Citing named authoritative sources | +28% AI visibility; 2.1x citation likelihood when claims are attributed | Forging authoritative sources to endorse yourself |
| Topical coverage across a subject | Coverage correlates 0.51 with AIO citation vs 0.09 for Domain Rating | Template-spun volume with no new information (answer hegemony / pseudo-consensus) |
| Clear, fluent, well-structured writing | +29% AI visibility (fluency optimisation) | Keyword stuffing — not banned by regulators, but it costs about 8% |
Note the one row with no regulator attached. Keyword stuffing is legal and still self-harming: the Princeton benchmark measured it at roughly −8% AI visibility. Compliance is a floor, not a strategy — you can be fully compliant and still invisible. The tactics that survive both tests are the ones built on things that can be checked.
How to audit a GEO vendor: six checks
- 1.Ask which of the four prohibited practices they rule out in writing
Corpus poisoning, answer hegemony, pseudo-consensus manufacturing, prompt injection. A vendor that cannot name all four has not read the standard it is selling against.
- 2.Ask which source tier the placements land in — and demand URLs, not counts
A-tier content is cited 7.2x more than C-tier. If a proposal is priced per article with no named publications, it is a volume play, and volume plays are what the platforms just started discounting.
- 3.Reject any absolute-effect promise
"Guaranteed first position", "secured placement", "locked into the category top three" are now prohibited claims, and are the single most reliable tell that the method depends on manipulation.
- 4.Require the measurement method to name its sample size
The September 2026 measurement standard attaches minimum sample sizes to its nine core metrics for a reason. SE Ranking measured only 9.2% URL overlap between repeated AI Mode runs, with 21.2% of queries returning zero overlap — a vendor reporting single-run screenshots is not measuring anything.
- 5.Require retained drafts and raw logs
Traceability is now an explicit compliance requirement: working drafts and run logs must be available for inspection, not sealed. Ask what you would receive if a regulator asked tomorrow.
- 6.Confirm the work is in-house
Subcontracting to unnamed studios is prohibited under the vendor framework, and only about 19% of the 200+ agencies selling GEO have genuine in-house technology. Ask who writes, who publishes, and under what byline.
What changed in September 2026
- ▸ Five group standards, 14 September 2026. Compiled by 41 organizations and 75 experts across terminology, vendor evaluation, pricing and measurement, trusted corpus, and compliance and security. Seven pricing units and five pricing models replace quote-only proposals; nine core metrics each carry a stated statistical calibre and minimum sample size.
- ▸ A one-vote veto on vendors. Providers with black-hat operations, false-effect claims, or data-security penalties are excluded outright — with a five-dimension evaluation framework covering qualifications, technology and data, delivery, quality assurance and sector practice.
- ▸ Mandatory human review. AI output must be human-verified, and finance and medical content requires 100% human review — the strictest tier, aimed squarely at YMYL categories where our own data shows AI Overviews trigger on 48.75% of searches in Health Care.
- ▸ Counter-GEO-Bench, 2 September 2026. An academic benchmark (accepted to EMNLP 2026) showing a brand cited beside misinformation it never published, with tested safety filters cutting attack success by only 5.7%.
- ▸ Platform-side countermeasures. DeepSeek down-weighting recent homogeneous marketing listicles; Kimi rolling citations back to pre-2025 sources. The defences are now on both sides of the retrieval step.
Frequently asked questions
What is AI corpus poisoning in GEO?
AI corpus poisoning is the practice of manipulating what a generative engine retrieves and repeats by flooding the web with fabricated or misleading material: fake specifications injected into low-authority sites, invented case studies, forged credentials, template-spun listicles, and fabricated expert endorsements. It differs from classic SEO spam in its target — the goal is not a ranking, it is to change what the model says about a brand inside a generated answer.
What GEO tactics are now explicitly banned?
T/CAPT 026-2026, published on 11 August 2026, names four prohibited practices: corpus poisoning, answer hegemony, pseudo-consensus manufacturing, and prompt injection attacks. The five group standards published on 14 September 2026 add commercial red lines, banning absolute-effect promises such as guaranteed first position, requiring human verification of AI output, and requiring 100% human review in finance and medical categories.
Does corpus poisoning actually work?
Partly, and that is the problem. Counter-GEO-Bench (arXiv, 2 September 2026, accepted to EMNLP 2026) showed a brand can be cited alongside misinformation it never published, and that tested safety filters cut attack success by only 5.7%. Retrieval is the soft target: Trellner found 59.8% of Perplexity citations come from domains ranked worse than #100,000. The defence is that platforms are responding — DeepSeek now down-weights recent homogeneous marketing listicles and Kimi has rolled citations back to pre-2025 sources.
What is the A/B/C/D source tier system?
T/CAPT 026-2026 grades sources into four credibility tiers. A covers government bodies, national databases and original state-media reporting; B covers academic papers, white papers and original mainstream reporting; C covers corporate sites and commercial reports, usable only as supporting reference; D covers anonymous sources and low-quality aggregators, which AI systems are explicitly told not to rely on. Monitoring cited A-tier content at 7.2x the rate of C-tier and more than 23x D-tier.
Can a GEO agency guarantee citations or a number one position?
No, and an agency offering one is now a compliance risk rather than a strong performer. The September 2026 group standards prohibit absolute-effect language because generative answers are assembled in real time from sources no vendor controls. The evaluation framework also sets a one-vote veto: providers with black-hat operations, false-effect claims or data-security penalties are excluded outright.
Is the brand liable if the agency did the poisoning?
Yes. Under the Internet Advertising Measures the advertiser is responsible for the truthfulness of the content, and that responsibility is not discharged by delegating to a third party. Once cross-verification exposes a fabricated endorsement, the penalty extends past the individual page to the entity record itself — name, address and organisation details get marked unreliable, and recovery costs far exceed the campaign.
How do I audit a GEO vendor?
Six checks: ask which of the four prohibited practices the vendor rules out in writing; ask which source tier its placements land in and demand named URLs rather than volume counts; reject any absolute-effect promise; require the measurement method to name its sample size and prompt panel; require drafts and raw logs to be retained for traceability; and confirm the work is performed in-house rather than subcontracted to unnamed studios.
How big is the GEO market in 2026?
China is the reference market because it is the only one publishing standards. iResearch sized the 2025 domestic GEO market at ¥12.7–28 billion, growing about 120% year over year; Analysys projects ¥4.8 billion in 2026 → ¥15.4 billion in 2027 → ¥43 billion in 2028, a CAGR above 130%. CNNIC counted 602 million generative-AI users at 42.8% penetration. More than 200 agencies sell GEO services, but only about 19% have genuine in-house technology.
Is legitimate GEO at risk of being caught by these rules?
Largely no, because the boundary is drawn on intent and verifiability rather than technique. Publishing original statistics, adding expert quotations and citing named sources remain legitimate and are measured at +33%, +41% and +28% in the Princeton KDD 2024 benchmark. What crosses the line is fabricating the statistic, inventing the expert, or manufacturing agreement across throwaway domains. Keyword stuffing is the edge case: not banned by regulators, but it costs about 8% of AI visibility.
What should a brand do in the next 30 days?
Run a six-point inventory: list every domain publishing claims about your brand and check whether you commissioned it; verify every statistic and endorsement on your own pages still resolves to a real source; remove template-spun pages; confirm your entity record is consistent across your site, Wikidata and major profiles; re-paper vendor contracts to remove absolute-effect language; and move part of the publishing budget into A-tier and B-tier sources, cited 7.2x more often than corporate pages.
Related GEO guides
References:
1. China Advertising Association (CAAC), five GEO group standards published 14 September 2026 — 41 drafting organizations, 75 experts; terminology, vendor evaluation, pricing and measurement, trusted corpus, compliance and security; seven pricing units, five pricing models, nine core metrics with minimum sample sizes; absolute-effect promises prohibited; five-dimension vendor evaluation with one-vote veto; 100% human review for finance and medical.
2. T/CAPT 026-2026, "Generative Engine Optimization (GEO) Trusted Information Dissemination and Information Ecosystem Governance Specification," published 11 August 2026 by the China News Technology Workers Federation, led by the Xinhua Net converged-media future research institute with Fudan University, Beijing University of Posts and Telecommunications, Shanghai Jiao Tong University and 30+ institutions — four prohibited practices and the A/B/C/D source-tier system.
3. Cyberspace Administration of China, "Qinglang" special campaign on AI application abuses (April 2026) — GEO malicious marketing classified as AI data poisoning.
4. Five-department rule on multi-channel internet information content distribution services, effective 1 September 2026 — 11 red lines; distribution duties raised from self-discipline to legal obligation.
5. Counter-GEO-Bench, arXiv preprint, 2 September 2026, accepted to EMNLP 2026 — brand cited alongside misinformation it never published; tested safety filters reduced attack success by 5.7%.
6. Aggarwal, P., Dugan, L., et al. "GEO: Generative Engine Optimization," arXiv:2311.09735, KDD 2024 — +33% statistics, +41% quotations, +28% citations, +29% fluency, −8% keyword stuffing.
7. CNNIC 57th Statistical Report on Internet Development in China — 602 million generative-AI users, 42.8% penetration.
8. iResearch and Analysys GEO market sizing, 2025–2028 — ¥12.7–28B (2025, +120% YoY); ¥4.8B (2026) to ¥43B (2028), CAGR above 130%.
9. Conductor 2026 AEO/GEO Benchmarks — 167,867,680 AI Overview citations; position one cited 24.9%, position ten 10.2%, 57% of citations outside the top 10; AI Overviews trigger on 48.75% of Health Care searches.
10. SE Ranking AI Mode repeat-run study (10,000 queries, 2026) — 9.2% mean URL overlap, 21.2% of queries with zero overlap. · Wellows cross-engine citation study (September 2026) — 10.2% URL overlap vs 67.4% brand overlap across 596,723 prompts. · Trellner, Perplexity citation rank analysis — 59.8% of citations from domains ranked worse than #100,000. · Authoritas GEO Signal Study — original statistics +156% AIO citation probability.
Want to check your site's GEO readiness?
Run the 27-point GEO auditRelated articles
9 Proven GEO Optimization Strategies (With Quantified Data)
Expert quotations boost AI visibility by 41%, statistics by 33%, fluency by 29%, citations by 28% — and statistics + citations compound to ~+61%. Updated August 2026 with validation from Conductor, GrackerAI, BrightEdge, Authoritas, and Previsible, plus the GEO market now worth $7.3B at 34% CAGR. Updated Aug 2026: ChatGPT passed 1B MAU and 32% of marketing leaders rank GEO their top 2026 priority (BrightEdge). The complete peer-reviewed guide to all 9 GEO strategies (Princeton, KDD 2024) with quantified lift percentages.
How to Get Cited by AI Search: The 2026 Playbook
Getting cited by ChatGPT Search, Perplexity, Google AI Overviews, Gemini, and Claude is a repeatable process, not luck. Updated September 2026: Ahrefs analysed 3 million+ US queries and found YouTube takes 22.9% of AI Overview citations (Reddit 18.5%, Facebook 10.1%), while LLM Pulse found 64.7% of citations on non-branded shopping queries go to brand and manufacturer sites. Also new this month: Google AI Mode entity-wrapped citations appear on 87.6% of recommendation prompts but 0% of informational ones, and a September 7 2026 replication found volume-controlled quotation and statistics edits produced no pooled lift. This 2026 playbook covers the 7-step citation workflow, the off-site reality that earned media drives 84% of AI citations (Muck Rack, May 2026), the 3.5B+ weekly AI queries you compete for, and the structural fixes (AI crawler access, Schema.org, FAQ) that move pages into AI answers.
AI Search Optimization: The Complete 2026 Guide
Updated August 2026: AI search optimization is the practice of structuring content so ChatGPT Search, Perplexity, Google AI Overviews, Gemini, and Claude cite it. AI search traffic grew 16× from 2024 to 2026, yet only ~12% of AI citations match Google top 10. This 2026 playbook covers the 9 GEO strategies with measured lift (+41% quotations, +33% statistics, +28% citations, +29% fluency), the 7-step optimization process, common mistakes (keyword stuffing −8%), and how to measure AI visibility.