Skip to main content
Back to Blog
12 min readBrass-SEO Team

Why ChatGPT Cites Some Pages and Skips Others

Two pages on the same topic. Similar word count. Similar Google rankings. ChatGPT cites one and ignores the other. Why?

The research now answers that question with specifics. Princeton's Generative Engine Optimization paper (Aggarwal et al., 2024) and Adam Gnuse's practitioner analysis of 7,500 ChatGPT referral sessions (Search Engine Land, 2025) converge on a small set of page-level traits that separate cited content from skipped content. Brass-SEO has documented seven. Most of them aren't what conventional SEO advice would predict.

What follows is a synthesis post, not a first-party study. Every claim cites a named source you can verify.

Quick Navigation


The Seven Traits That Separate Cited From Skipped

Brass-SEO's review of the GEO research literature surfaces seven page-level traits that the data consistently associates with higher ChatGPT citation rates. Pages carrying most of them get cited. Pages missing most of them get skipped, regardless of Google ranking.

# Trait Primary research source Effect on citation
1 Self-contained answer capsule after each heading Gnuse 2025 Present in 72.4% of cited posts
2 Links removed from capsule sentences Gnuse 2025 91% of cited capsules contained no links
3 Named expert quotes Princeton 2024 +41% citation rate (largest single-factor lift)
4 Original data and specific numbers Princeton 2024 + Gnuse 2025 +30–40% (statistics) / 52.2% of cited posts (original data)
5 Substantive content beyond fluency edits Princeton 2024 Fluency-only improvements had negligible effect
6 No keyword stuffing Princeton 2024 Keyword stuffing actively reduced citations
7 Brand name (not "we" or "our") in capsules Brass-SEO documented pattern Citation context strips surrounding text; brand must travel with the quote

The pattern is consistent. AI citation is driven by substance and structure, not by length, polish, or keyword density. The next seven sections cover each trait in detail.


Trait 1: A Self-Contained Answer Capsule

A page is more likely to get cited when each H2 is immediately followed by a one- or two-sentence statement that fully answers the heading's implied question. Brass-SEO calls these answer capsules. The Gnuse 2025 research found them in 72.4% of pages ChatGPT cited.

The Gnuse study, published on Search Engine Land, manually audited pages receiving 7,500 ChatGPT referral sessions across 15 domains. The strongest structural pattern was an extractable passage placed immediately after a heading. ChatGPT appears to lift these passages directly. Pages with extractable passages get cited. Pages that bury the answer don't.

A self-contained capsule passes two tests:

  1. Structure test: It makes sense as a standalone statement without reading the heading above it.
  2. Substance test: It contains something worth citing — a specific number, method, or insight that isn't common knowledge.

Generic statements like "page speed matters for SEO" fail the substance test even when they pass structure. Specific statements like "Google's Core Web Vitals threshold for Largest Contentful Paint is 2.5 seconds, above which pages see measurable ranking drops in mobile search results" pass both.

For the deep-dive on this trait, see Part 1: Answer Capsules: The Content Trait LLMs Cite Most.


Citation rates climb when the first one or two sentences after a heading carry no embedded links. Gnuse 2025 found that 91% of capsules ChatGPT cited had no links inside them.

The mechanism is concrete. When an AI system extracts a passage for citation, it strips the surrounding HTML. Hyperlinks vanish. A capsule that depends on a link for meaning ("see [this study]") becomes ambiguous after extraction. AI systems can recognize this fragility and prefer link-free passages that hold their meaning when isolated.

This doesn't mean removing all links from a post. It means moving them out of the capsule sentences into the supporting paragraphs. The capsule is the extractable summary. The body is where citations live.

The practical pattern:

## Heading

[Capsule sentence with no links — citable as-is.]
[Optional second capsule sentence — still no links.]

The supporting body paragraphs go here, and these can contain
all the [linked references] and [outbound citations] you want.

Brass-SEO formats every blog post this way after adopting the finding.


Trait 3: Named Expert Quotes

A page becomes more likely to be cited by AI systems when it carries direct quotations from named, recognized experts. The Princeton GEO study (Aggarwal et al., 2024) measured a +41% citation lift from this single change. It was the largest effect of any strategy the researchers tested.

The team compared baseline pages against versions modified by adding expert quotations. Nothing else changed: same structure, same length, same keyword density. The pages with added quotes were cited 41% more often across multiple AI systems.

The likely mechanism: expert quotes signal that a page is a primary or secondary source rather than a generic restatement. AI citation behavior rewards source-like content because it reduces the risk of citing thin or derivative material.

The bar is lower than most writers assume. "Expert" in this research context means a named, identifiable authority. A quoted researcher, founder, working practitioner, or organization spokesperson all qualify. What doesn't qualify: fabricated sources, anonymous attributions ("experts say"), or paraphrases ("according to research").

For the detailed framework on what counts, see Why Expert Quotes Boost AI Citations by 41%.


Trait 4: Original Data and Specific Numbers

Pages with original data get cited more often than pages built entirely from secondary references. Gnuse 2025 found original data in 52.2% of ChatGPT-cited blog posts. The Princeton study measured a +30–40% lift from adding statistics and a similar +30–40% lift from adding source citations.

Original data doesn't require a Princeton-scale study. It can mean:

  • "Brass-SEO crawled 487 pages from one customer's site and found a median title tag length of 58 characters."
  • "The five blog posts we tested with answer capsules saw a 23% impressions increase in the following 30 days versus the prior 30."
  • "Our payment processor records a 3-day median trial-to-paid conversion window."

These are first-party observations, attributable to the publisher, verifiable in principle. AI systems treat that kind of content as primary source material.

When original data isn't available, the next-best move is a specific cited statistic from a named source with a date. "According to the Princeton GEO study (Aggarwal et al., 2024), expert quotes produced a +41% citation lift" beats "research shows expert quotes help" every time.


Trait 5: Substance Over Fluency

Polishing the writing without changing the underlying claims doesn't raise citation rates. The Princeton 2024 study tested fluency optimization — rewriting the same content in more polished prose without adding new information — and measured a negligible effect on AI citations.

This is the most counterintuitive finding in the research. Common content advice prioritizes readability and tone. Those qualities matter for human readers. They don't move the needle on AI citation. What moves the needle is what the content says, not how nicely it says it.

The practical implication: time spent on substance produces more citation lift than time spent on polish. A well-written generic paragraph won't beat a clunky paragraph carrying a fresh statistic.

This isn't a license to publish badly written content. Fluency still matters for Google ranking, social sharing, and the reader experience. For AI citation specifically, the research is clear. Substance is the variable. Fluency is not.


Trait 6: No Keyword Stuffing

Pages with high keyword density get cited less often by AI systems, not more. The Princeton 2024 study measured a negative effect from keyword stuffing. It was the only modification of the nine tested that reduced citations.

The mechanism is direct. AI systems are trained to recognize low-quality content patterns. Unnatural keyword repetition is one of the strongest such signals. A page repeating "best SEO tool for small business" twelve times in 800 words triggers exactly the heuristic the model is built to penalize.

This finding overlaps with what Google has penalized for over a decade. The GEO research confirms the same penalty applies to AI citation behavior. Possibly more strongly, because AI systems aren't constrained by ranking-signal commitments and can simply skip the page.

The practical rule: write for the question, not for the keyword. A primary keyword appearing naturally 3–6 times in a 2,000-word post is sufficient. If you find yourself forcing the keyword into sentences where it doesn't fit, you're past the threshold.


Trait 7: Brand Name in the Capsule

When AI systems cite a capsule, they extract the passage but not always the surrounding context that identifies the publisher. A capsule that says "we connect to Google Search Console" gives the reader no way to know who "we" is. A capsule that says "Brass-SEO connects to Google Search Console" carries its attribution with the citation.

The public research doesn't measure this directly. It follows from the same extraction mechanism Gnuse 2025 documented. The cited passage travels. Everything else can be stripped. Brass-SEO documents this pattern internally and applies it across every post.

The pattern applies to product references too. "Our AI Citability button" becomes "the Brass-SEO AI Citability button." A reader landing on a citation knows what product is referenced.

The same principle works for any publisher. Replace "we," "our," "our tool," and "our service" with your brand name inside any capsule sentence. The rest of the post can use "we" freely. Only the first one or two sentences after each H2 need the brand-attribution treatment.


How to Audit Your Pages Against These Seven Traits

Run this checklist against each post you want AI to cite. The Brass-SEO AI Citability page audit automates the same checks through the dashboard. The manual version takes about three minutes per page.

  1. Pick an H2 section. Read the first two sentences immediately after it.
  2. Structure test: Does that passage make sense if you read it without the heading?
  3. Substance test: Does it contain a specific fact, number, or claim worth citing?
  4. Link check: Any hyperlinks in the first two sentences? Move them to the supporting paragraphs.
  5. Quote check: Does the post quote at least one named expert with attribution?
  6. Data check: Does the post contain at least one specific number, statistic, or measurement?
  7. Density check: Does any keyword appear more than once every 200 words? If yes, dilute.
  8. Brand check: Does the capsule say "we" or "our" instead of your brand name? Replace.

A page that passes all eight checks on every H2 is structurally optimized for AI citation. That isn't the same as guaranteed citation. The topic still has to be one AI systems are asked about. But it removes the structural barriers the research has shown to suppress citations.

For the broader GEO framework that contextualizes these traits, see the Generative Engine Optimization guide. For how content length interacts with these traits, see Content Length and AI Citations: What Data Shows.


Frequently Asked Questions

Does Google ranking predict ChatGPT citations?

Not reliably. Gnuse 2025 found pages with strong Google rankings that received zero ChatGPT citations, and pages with modest rankings that received many. The traits that drive AI citation overlap partially with Google's ranking signals. They aren't the same. A page can rank well and still be skipped by AI. A page can rank modestly and still be cited if it carries the right structural traits.

How many of the seven traits does a page need?

The research doesn't give a strict cutoff. The pattern is additive. Pages carrying more traits get cited more often. Brass-SEO's recommendation: treat traits 1, 2, and 3 (capsules, link-free capsules, expert quotes) as the highest-priority three. These have the strongest research support. Traits 4–7 are reinforcing.

What about other AI systems besides ChatGPT?

The Princeton 2024 study tested multiple AI systems and found the +41% expert-quote effect held across them. The Gnuse 2025 study focused on ChatGPT referral data, so the structural findings (capsules, link-free zones) are best-validated for ChatGPT. Anecdotal data from Perplexity, Claude, and Gemini suggests the same patterns apply with less rigorous validation.

How long should a capsule be?

The research range is 120–150 characters. Shorter capsules can work but tend toward generic, failing the substance test. Longer capsules approach paragraph length and lose extractability. The 120–150 character window aligns with how AI systems chunk text for citation.

Is original data really necessary?

It's one of the strongest traits. It isn't the only path. Pages without original data still get cited when they carry strong capsules, expert quotes, and specific cited statistics from other sources. Original data is a force multiplier on the other traits, not a prerequisite.

Where can I read the source research directly?

The Princeton GEO study is Aggarwal et al., "GEO: Generative Engine Optimization" (2024). The Gnuse practitioner research was published on Search Engine Land in 2025 under the title "How to get cited by ChatGPT: The content traits LLMs quote most." Both are publicly accessible.

Ready to try Brass-SEO?

Get AI-powered SEO insights from your Google Search Console and Analytics data.