Skip to main content
Back to Blog
11 min readBrass-SEO Team

AI Citation Research: What the Papers Say

Search "how to get cited by ChatGPT" and you will find hundreds of posts confidently listing rules: add statistics, write shorter paragraphs, use bullet points. Almost none of them link to a study. The advice sounds reasonable, but reasonable is not the same as tested.

Copper Sun Content and Creative, LLC — the same company that builds Brass-SEO — runs a separate site, coppersun.io, for AI marketing tools. That site maintains a Research Center: a curated page of peer-reviewed and preprint papers on how generative AI systems select and cite sources, plus a blog post walking through what those papers found. This post covers what is actually in both, disclosed plainly and cited accurately, so you have a sourced starting point instead of another unsourced list.

Quick Navigation


Full Disclosure: Same Company, Different Product

Copper Sun Content and Creative, LLC builds a small family of independent AI tools, and Brass-SEO is one of them. Coppersun.io is another product from the same company, built for marketing teams running full campaigns rather than solo business owners checking their own SEO. We've written before about the full lineup, including why the products stay separate rather than bundled.

That relationship is disclosed here because it matters. Brass-SEO is not an independent third party evaluating coppersun.io's research page. It's a sister product pointing at work the same company published. The research itself is still real and the citations below are accurate, but you should know who is telling you that before you decide whether to trust it.

The Four Papers in the Research Center

Coppersun.io's AI Citation Research Center functions as an annotated literature review, not a raw bibliography. Each entry names the paper, the authors, the venue, and a short explanation of how the finding connects to content strategy. Four papers are listed on the Research Center page:

Paper Authors Venue Focus
GEO: Generative Engine Optimization Aggarwal et al. ACM KDD 2024 Tests nine content strategies for AI citation impact
Citation Selection vs. Citation Absorption in AI Search Zhang et al. arXiv preprint 2604.25707 (2026) Separates retrieval from appearing in the generated answer
Structural Feature Engineering for Generative Engine Optimization Yu et al. arXiv preprint 2603.29979 (2026) Tests document architecture and formatting independent of content
News Source Citing Patterns in AI Search Systems Yang arXiv preprint 2507.05301 (2025) Measures how citation authority concentrates over time

One of these four, the Aggarwal paper, also shows up in Brass-SEO's own GEO research roundup, which covers three separate research streams. The other three papers here — Zhang et al., Yu et al., and Yang — are not covered in that post. This piece exists to fill that gap.

Three of the four papers are 2025 or 2026 arXiv preprints, not journal articles. That's not a red flag by itself. Preprints let a fast-moving field share findings before the slower peer-review cycle catches up, and computer science research routinely circulates this way before formal publication. It does mean you should weigh a preprint's claims a notch below a peer-reviewed paper's, and the FAQ below explains that distinction in more detail.

What Aggarwal et al. Actually Tested

Aggarwal et al.'s "GEO: Generative Engine Optimization" (ACM KDD 2024) tested nine content modification strategies across five different generative engines: Bing, You.com, Perplexity, NeevaAI, and Llama. The paper is the source of the field's name.

The headline result: adding concrete numerical data to a page raised source visibility in AI-generated answers by up to 40% in controlled testing. That single finding gets cited constantly in GEO advice, often without attribution. Knowing it comes from a peer-reviewed, multi-engine study changes how much weight it deserves compared to a blog post asserting the same number with no source.

Testing across five separate engines instead of one is what makes the result worth taking seriously. A finding that only holds up on a single AI system could just as easily be an artifact of that one system's particular retrieval quirks. A finding that holds across Bing, You.com, Perplexity, NeevaAI, and Llama is describing something closer to a general property of how these systems select sources, not a one-off preference for a single platform's algorithm.

Selection Versus Absorption: The Zhang et al. Distinction

Zhang et al.'s 2026 preprint (arXiv 2604.25707) makes a distinction most GEO advice skips entirely: getting selected as a source and getting absorbed into the generated answer are two different problems.

A page can be retrieved by an AI system's search step and still contribute nothing to the final response if the system doesn't judge the content worth quoting or paraphrasing. The paper identifies four traits associated with high absorption rates, separate from whatever gets a page retrieved in the first place. This matters for anyone treating "getting indexed by an AI system" and "getting cited by it" as the same goal. According to Zhang et al., they aren't.

Think about the practical version of this. Perplexity might pull your page into its candidate results for a query, technically making you part of the retrieval set, and then quote a competitor's page instead because yours didn't offer a passage clean enough to lift into the answer. You'd never see that near-miss in any dashboard. Traditional SEO tools measure whether you rank. Nothing measures whether you were considered and passed over, which is exactly the gap Zhang et al.'s distinction points at.

Structure as Its Own Variable: Yu et al.

Yu et al.'s 2026 preprint (arXiv 2603.29979) isolates document architecture, information organization, and visual formatting as variables independent of the underlying content quality. The paper's finding: structural optimization measurably improves citation likelihood on its own. Good writing alone doesn't do the same work.

That distinction, structure as a standalone lever rather than a side effect of good writing, lines up with a separate finding already documented on Brass-SEO's blog: the Princeton research behind Aggarwal et al. found that fluency improvements alone had negligible citation impact, while structural and substantive changes did. Two different research groups, working from different angles, reaching a compatible conclusion about structure mattering on its own terms.

Why Citation Authority Concentrates: Yang

Yang's 2025 preprint (arXiv 2507.05301) studies news source citing patterns across AI search systems and finds that citation authority concentrates among a small set of sources, and that concentration is stable rather than resetting each cycle. Once a source establishes itself as a system's preferred citation for a topic, that position tends to persist.

The practical implication is less encouraging than the other three papers: getting cited once does not guarantee an even playing field going forward. Established sources keep their position. That's a reason to treat GEO as an ongoing effort rather than a one-time fix, not a reason to skip it.

For a small business, that finding cuts two ways. The discouraging read is that a handful of large publishers or well-established sites may already hold the citation slots for your topic, and dislodging them takes sustained work rather than a single optimized page. The more useful read is that concentration Yang describes is measured at the topic level, not the whole internet at once. A large competitor's authority on their core keywords says nothing about who holds the citation slot for the specific, narrow question your business actually answers well. Yang's paper studied news sources specifically, and how cleanly that concentration effect generalizes to smaller commercial sites and local business topics is still an open question the paper doesn't answer.

What the Coppersun.io Blog Post Argues

Coppersun.io's companion blog post, "The Research Behind AI Citation: What the GEO Papers Actually Found," walks through the same four papers and draws four conclusions from them: concrete numerical data measurably improves visibility, structural formatting matters as its own variable rather than a side effect of quality writing, citation selection and citation absorption are separate problems that need separate fixes, and citation authority concentrates among established sources over time rather than resetting.

The post also names its own limits. Four papers is a real but still small evidence base for a field this new, and the post says so rather than overstating what the research settles. You can read the full post here. That's the same posture this post is taking: name what the papers found, name what they don't cover, and don't round up.

Why Sourced Research Beats Another Opinion Post

Most GEO content on the web is speculation dressed as expertise. A post says "AI systems prefer well-structured content" with nothing behind it beyond the author's own observation of what seemed to work on their site. That might be true. It also might be a pattern the author noticed in a handful of cases and generalized past what the evidence supports.

The difference between that and a cited paper is falsifiability. Aggarwal et al.'s claim that statistics addition improves visibility by up to 40% was measured across five AI engines with a controlled test design, published through ACM KDD, and can be checked by anyone willing to read the paper. A blog post's claim that "adding numbers helps" cannot be checked at all, because there's nothing underneath it to check.

This doesn't mean every unsourced GEO claim is wrong. It means you have no way to tell which unsourced claims are right without testing them yourself. A bibliography like coppersun.io's Research Center gives you a starting point that other people have already scrutinized, which is a different thing than a starting point one blogger invented and everyone else copied.

Where This Leaves Small Business Owners

You don't need to read four academic papers to run your SEO. What's useful here is narrower: when you read a GEO claim anywhere, including on this blog, ask whether it points to something you could go check. "Add expert quotes" backed by Aggarwal et al.'s 41% finding is a different kind of claim than "add expert quotes" backed by nothing.

Brass-SEO's own AI Citability audit is built on the research documented in Brass-SEO's GEO research roundup — the Aggarwal paper covered above, plus the practitioner and platform-level findings that post covers in full. If you want to see how your own pages score against that framework, connect Brass-SEO to your Google Search Console and GA4 from your dashboard and ask the AI to run an AI Citability check on any page. For the broader strategy behind all of this, the Generative Engine Optimization guide covers the full framework this research supports.


Frequently Asked Questions

Is coppersun.io a competitor to Brass-SEO?

No. Both are built by Copper Sun Content and Creative, LLC. Coppersun.io is priced for marketing teams running full campaigns, and Brass-SEO is priced for individual business owners managing their own SEO. We've covered the full product family and how the products differ.

Are these four papers peer-reviewed?

One is. Aggarwal et al.'s "GEO: Generative Engine Optimization" was published through ACM KDD 2024, a peer-reviewed venue. The other three (Zhang et al., Yu et al., and Yang) are arXiv preprints, meaning they've been posted publicly but have not necessarily completed formal peer review. That distinction matters when you're weighing how much confidence to put in each finding.

Does Brass-SEO's own GEO research post cover the same papers?

Mostly not. Brass-SEO's GEO research roundup covers the Aggarwal et al. Princeton study in depth, along with separate practitioner research (Adam Gnuse's answer capsule analysis) and platform-level citation data. The Zhang et al., Yu et al., and Yang papers covered in this post are not discussed there.

What does "citation selection versus citation absorption" mean in plain terms?

Selection is whether an AI system's search step retrieves your page as a candidate source at all. Absorption is whether the system actually quotes or paraphrases your page in the answer it gives the user. Zhang et al.'s 2026 preprint argues these are separate problems with separate fixes, so a page can be retrieved constantly and still contribute nothing to what users actually read.

Should I read the actual papers myself?

If you want to verify any claim in this post, yes. Aggarwal et al. is indexed at arXiv (search "GEO Generative Engine Optimization Aggarwal 2024"), and the other three are searchable by their arXiv preprint numbers listed above. Reading the source is always the most reliable way to confirm a claim someone else summarized for you, including this post.

Ready to try Brass-SEO?

Get AI-powered SEO insights from your Google Search Console and Analytics data.