Skip to main content

AI Citation & Generative Engine Optimization

Brass-SEO · 6 entries · last verified August 2026

Brass-SEO tracks the primary empirical literature on generative engine optimization — the emerging field that asks why some sources appear in AI-generated answers and others do not. The studies below represent the peer-reviewed core of what is currently known.

Contents — 6 entries
  1. 1.GEO: Generative Engine Optimization
  2. 2.Citation Selection vs. Citation Absorption in AI Search
  3. 3.Structural Feature Engineering for Generative Engine Optimization
  4. 4.News Source Citing Patterns in AI Search Systems
  5. 5.Source Coverage and Citation Bias: LLM Search vs. Traditional Search
  6. 6.GEO-16: An Empirical Signal Framework for Generative Engine Optimization
  7. Frequently Asked Questions

GEO: Generative Engine Optimization

Aggarwal et al., 2024. ACM KDD 2024.

Brass-SEO draws on this to establish the foundational evidence for AI search optimization. Aggarwal et al. tested nine content strategies across five generative engines and found that Statistics Addition — adding concrete numerical data to a page — raised source visibility by up to 40% in their test matrix, while adding citations and quotations also improved visibility. Keyword stuffing and fluency edits provided no measurable benefit. Content with verifiable numerical claims and sourced references is what generative engines lift; content that merely reads well does not get a corresponding benefit.

Examines:
The first rigorous study of which content modifications increase a source's likelihood of appearing in generative engine answers, tested across Bing, You, Perplexity, NeevaAI, and Llama.
Brass-SEO draws on:
The 40% visibility gain from Statistics Addition — the basis for Brass-SEO's recommendation to add concrete numerical evidence to pages targeting AI citation.

Citation Selection vs. Citation Absorption in AI Search

Zhang et al., 2026. arXiv preprint 2604.25707.

Brass-SEO uses this as the framework for distinguishing two separate problems: getting selected (does the engine pull your URL into the context?) versus getting absorbed (does your page's content actually shape the generated answer?). Zhang et al. identified four traits of high-absorption pages — greater length, structured formatting, semantic alignment with the query, and extractable numerical evidence. A page can pass retrieval and still contribute nothing to the final answer. Brass-SEO's AI Citability analysis surfaces both failure modes, not just indexability.

Examines:
Measurement framework separating citation selection from citation absorption across AI search platforms, with analysis of which page characteristics predict absorption.
Brass-SEO draws on:
The four high-absorption traits — the basis for Brass-SEO's content structure recommendations in AI Citability reports.

Structural Feature Engineering for Generative Engine Optimization

Yu et al., 2026. arXiv preprint 2603.29979.

Brass-SEO draws on this to quantify what structural formatting delivers independent of semantic content. Yu et al. tested optimization across three structural levels — document architecture, information organization, and visual emphasis via headers, lists, and callouts — and measured a statistically significant improvement in both citation rate and citation quality across all three levels in their experiments. Restructuring an existing page into a scannable, hierarchical format produces measurable AI citation improvement without rewriting the underlying information.

Examines:
Three-level structural optimization framework (document architecture, information organization, visual emphasis) and its measured effect on AI citation rate and quality.
Brass-SEO draws on:
The finding that structure-level changes improve citation independently of content — cited when explaining why page formatting appears in Brass-SEO's AI Citability analysis.

News Source Citing Patterns in AI Search Systems

Yang, 2025. arXiv preprint 2507.05301.

Brass-SEO cites this when explaining citation concentration in AI search. Yang analyzed over 366,000 citations across 24,000 conversations from ChatGPT, Perplexity, and Google AI Overviews and found that AI search systems cite a small and skewed set of established sources regardless of query type. The engines function as gatekeepers that amplify existing authority hierarchies. A well-placed mention in a high-authority publication within a category is likely worth more AI citation value than optimizing a dozen of your own pages, because generative engines inherit and amplify the authority structures baked into their training data.

Examines:
Citing patterns of AI search systems across a large-scale multi-platform dataset covering ChatGPT, Perplexity, and Google AI Overviews.
Brass-SEO draws on:
The authority concentration finding — explains why Brass-SEO's GEO playbook includes off-site entity mentions alongside on-page optimization.

Source Coverage and Citation Bias: LLM Search vs. Traditional Search

Zhang et al., 2025. arXiv preprint 2512.09483.

Brass-SEO monitors this for evidence that AI and traditional search draw from meaningfully different source pools. Zhang et al. analyzed 55,936 queries and found that 37% of domains cited by LLM-based search engines do not appear in traditional search results for the same queries. Reference sites, technical documentation, and academic sources are surface categories that AI engines favor and Google frequently underranks. Practitioners who use Google rank as the sole proxy for AI citation visibility are, according to this analysis, missing more than a third of the opportunity surface.

Examines:
Comparison of domain and source coverage between LLM-based search engines and traditional search for identical queries across 55,936 examples.
Brass-SEO draws on:
The 37% unique-domain finding — the factual basis for Brass-SEO's argument that GEO requires a separate audit from traditional rank-tracking.

GEO-16: An Empirical Signal Framework for Generative Engine Optimization

Kumar & Palkhouski, 2025. arXiv preprint 2509.10762.

Brass-SEO uses this as an empirical signal checklist. Kumar and Palkhouski measured citations across Brave Search, Google AI Overviews, and Perplexity and found that metadata freshness, Semantic HTML structure, and schema.org markup showed the strongest statistical association with citation across all three engines in their dataset. Schema markup and kept-current metadata are the highest-leverage quick wins they identified — measurable, engine-agnostic, and available to any site without rewriting existing content.

Examines:
Empirical analysis of 16 GEO signals across three AI search platforms measuring which on-page signals correlate with citation frequency.
Brass-SEO draws on:
The metadata freshness, Semantic HTML, and schema.org signal findings — cited when prioritizing recommendations in Brass-SEO's AI Citability analysis.

Frequently Asked Questions

What is generative engine optimization (GEO)?

GEO is the practice of modifying web content to increase the likelihood that generative AI search engines — such as Bing Copilot, Perplexity, and Google AI Overviews — cite the page in their generated answers. The term and its initial framework were formalized in Aggarwal et al. (2024), which tested nine content modification strategies across five generative engines. GEO is distinct from traditional SEO, which targets ranking in ten blue links; GEO targets inclusion in a generated answer.

Which type of content change has the strongest evidence for increasing AI citations?

The 2024 Aggarwal et al. GEO study found that adding concrete numerical statistics to a page produced the largest improvement — up to 40% in their test matrix — followed by adding citations and quotations. Keyword stuffing and fluency improvements showed no measurable benefit. The 2025 GEO-16 study also found schema.org markup and metadata freshness to be strongly associated with citation across three AI search platforms.

Does ranking well in Google guarantee AI citation?

No. Zhang et al. (2025) analyzed 55,936 queries and found that 37% of domains cited by LLM-based search engines do not appear in traditional search results for the same queries. AI engines favor reference sources, technical documentation, and academic content that traditional search often underranks. Google rank is a useful signal but not a reliable proxy for AI citation eligibility.

What is the difference between citation selection and citation absorption?

Citation selection refers to whether a search engine pulls a URL into its retrieval context for a given query. Citation absorption refers to whether the page's content actually shapes the generated answer — whether the engine quotes, paraphrases, or cites it. Zhang et al. (2026) found these are separable: a page can pass retrieval and still contribute nothing to the final answer. High-absorption pages tend to be longer, structured, semantically aligned with the query, and rich in extractable numerical evidence.