Skip to main content
Back to Blog
9 min readBrass-SEO Team

How to Test Your Brand's AI Search Visibility

If you ask "are we showing up in AI search?" there's no AI-search equivalent of Google Search Console to check. The data isn't published anywhere.

You have to test it yourself. Query each AI system the way a customer would. Record what comes back. This post is the method Brass-SEO uses to do that — a repeatable quarterly process that takes about 90 minutes and gives you a real baseline.

The goal isn't perfection. AI systems produce different answers to the same query at different times. The goal is a directional measurement: which queries cite your brand, which cite competitors, which cite nobody, and whether that distribution changes quarter over quarter.

Quick Navigation


Why You Have to Test This Manually

Brass-SEO's AI search visibility method relies on manual querying because no AI vendor publishes per-domain visibility data. ChatGPT, Perplexity, Gemini, and Claude don't offer the equivalent of a Search Console for their citation patterns. The only reliable measurement is to run the queries yourself and record the results.

The Princeton GEO study (Aggarwal et al., 2024) and Adam Gnuse's 2025 practitioner research both relied on manual or programmatic querying for the same reason. Tools that claim to track AI visibility generally run automated queries on your behalf and report what comes back. That's the same thing this method does, just by hand.

The advantage of manual testing is that you see what real users see, including the wording variations and follow-up suggestions the AI presents. The disadvantage: it doesn't scale beyond 50–100 queries per session. That's why this method focuses on a curated query set tested quarterly, not a continuous monitor.


Step 1: Build Your Query List

The first quarterly run starts with a list of 20–40 queries covering three categories. Brass-SEO recommends roughly 50% category 1, 30% category 2, 20% category 3.

Category 1: Branded queries. "What is [your brand]?" "Is [your brand] worth it?" "[your brand] vs [competitor]." These should always cite you if AI knows your brand exists. If they don't, you have a brand-visibility problem.

Category 2: Solution queries. Queries where your brand is one right answer but not the only one. "Best [tool category] for [use case]," "alternatives to [competitor]," "how to [do the thing your product does]." These are the most strategically important. They reveal whether AI views you as a credible option.

Category 3: Topical queries. Industry-knowledge questions where your blog content is the type of source AI would cite. "What is [industry term]?" "How does [process] work?" These test whether your content marketing is reaching AI's citation set.

For a SaaS example, a category 2 query might be "best free alternative to Semrush." A category 3 query might be "what is generative engine optimization." For a local service business, the equivalents would be "best [service] in [city]" and "how to choose a [service provider]."

Write the queries in natural language. The way a customer would actually type them. Don't include your brand name in category 2 or 3 queries. That defeats the purpose.


Step 2: Run Each Query on Four Engines

Brass-SEO tests the same queries across four AI search engines: ChatGPT, Perplexity, Gemini, and Claude. Each engine has different training data, different real-time web access, and different citation patterns. Visibility on one doesn't predict visibility on the others.

Use a fresh session for each engine to avoid history bias. Most engines personalize follow-up questions based on prior interactions. The same query in an existing session can return different sources than in a new one. Either open a new chat or use private/incognito mode.

Run each query exactly as written. Don't iterate or refine. The goal is to capture what a first-time user would see, not what an experienced prompter could extract. If the engine asks for clarification, pick the most natural follow-up and continue.


Step 3: Record What You See

For each query on each engine, record four things in a spreadsheet:

  1. Whether your brand was cited: Yes / No / Mentioned but not cited (named in the answer but not in the source list)
  2. Where your brand appeared: Source list / inline citation / both / neither
  3. Which competitors were cited: A comma-separated list of competitor brands that appeared
  4. The URL that was cited if your brand was: Often the specific blog post or product page

A simple spreadsheet structure works: one row per query, one column per engine, doubled for the four data points above. That's 8 columns per engine, 32 columns total. Manageable in any spreadsheet.

Take a screenshot of any query that surprised you. Cited a page you didn't expect. Missed a page that should have been obvious. The screenshots are useful for the quarter-over-quarter comparison in Step 6.


Step 4: Score Visibility by Tier

After running all queries on all engines, score the results into three tiers per category:

  • High visibility: Cited on 3 or 4 of the 4 engines
  • Medium visibility: Cited on 1 or 2 of the 4 engines
  • No visibility: Cited on 0 engines

Calculate the percentage of queries in each category that hit each tier. A healthy distribution for an established brand looks roughly like:

Category High Medium None
Branded 90%+ 5–10% <5%
Solution 30–50% 30–40% 20–30%
Topical 10–25% 30–40% 40–60%

These are Brass-SEO's observed ranges from testing established SaaS brands. New or niche brands will skew lower. The point isn't to hit specific numbers. It's to get a baseline you can compare against in three months.


Step 5: Compare Against Competitors

The query data also reveals competitive position. For each category 2 query, count how many times each competitor was cited across the four engines. The competitor with the most citations is the AI-search incumbent for that query.

When a single competitor dominates 70%+ of category 2 queries, you have a visibility deficit on that query set. Your content strategy should target the same queries with capsule-optimized pages. For the structural traits that make pages more citable, see Why ChatGPT Cites Some Pages and Skips Others.

When citations are distributed across many competitors with no dominant player, the query set is open. Targeted content has a high chance of moving into the top citation set within a few quarters.


Step 6: Repeat Quarterly and Track Drift

The first run produces a baseline. The value comes from running the same query set again 90 days later and comparing.

Track three drift metrics quarter-over-quarter:

  1. Citation count change per category (how many more or fewer queries cited you)
  2. Tier movement (queries that moved from "no visibility" to "medium" or from "medium" to "high")
  3. Competitive churn (which competitors gained or lost citations)

A 10–20% citation improvement quarter-over-quarter is meaningful. Larger swings are common in the first few quarters as your content marketing finds its footing in AI training data. The method gets more valuable the longer you run it. By year two, you have a real time-series of AI visibility for your brand.

For the broader GEO framework that contextualizes this measurement, see the Generative Engine Optimization guide.


Common Mistakes and How to Avoid Them

Mistake 1: Testing with logged-in accounts that have history. Personalization skews results. Always use new sessions or incognito mode.

Mistake 2: Refining queries to get better answers. The point is to measure what cold queries return. Not what a skilled prompter can extract. Run the query as written and record the first response.

Mistake 3: Only testing branded queries. Branded queries should always work. They're a hygiene check, not a strategic signal. The real value lives in category 2 (solution queries) and category 3 (topical queries).

Mistake 4: Counting all engines as equal. ChatGPT and Perplexity have far more users than Claude and Gemini as of 2026. A citation on ChatGPT is worth more in raw traffic than one on Claude. Weight your scoring or report engines separately if traffic distribution matters.

Mistake 5: Quitting after one run. A single quarterly run is a baseline, not a signal. The method rewards consistency. Run it on the same date each quarter for a year before drawing conclusions.

For the related question of what content traits drive AI citations in the first place, see Why Expert Quotes Boost AI Citations by 41% and Part 1: Answer Capsules: The Content Trait LLMs Cite Most.


Frequently Asked Questions

How long does the full test take?

For 30 queries across 4 engines, expect 60–90 minutes of focused work. The query list and spreadsheet structure are reusable. Subsequent quarterly runs are faster — typically 45–60 minutes once you have the system set up.

Can this be automated?

Partially. Several commercial tools run automated AI visibility testing. They have limits: they can't replicate logged-in user experiences, they can't follow up the way a human can, and they often miss the "mentioned but not cited" distinction. Manual testing remains the most reliable signal as of 2026.

What if my brand has zero citations on the first run?

That's common for new or niche brands. It means your starting baseline is zero. Every quarter from here is a measurement of growth from that floor. Pair the method with a content strategy targeting your category 2 and category 3 queries. Capsule-optimized pages with named expert quotes are the highest-impact moves per the public research.

How do I pick the right competitors to test against?

Three approaches: competitors you know about from your industry, the brands that appeared in your own test results for category 2 queries, and the brands a customer interview would name as alternatives. Combine all three.

Does Google AI Overviews or AI Mode count as an engine?

Yes. As of 2026, Google AI Overviews and the broader AI Mode experience are significant AI search surfaces. Brass-SEO recommends testing them as separate engines from traditional Google search.

Ready to try Brass-SEO?

Get AI-powered SEO insights from your Google Search Console and Analytics data.