
AI visibility tracking means measuring whether AI answer engines mention and cite your brand when buyers ask questions in your category. Most tools report a mention count. That number alone is not actionable. In our own measurement of 23 buyer prompts, 91% triggered a Google AI Overview and we were cited in zero of the non-branded ones — and the reason turned out to be something a mention counter would never have shown us.
This is not a roundup of tools. We built a tracker, pointed it at our own category, and are publishing what came back, including the parts that make us look bad. If you are evaluating AI visibility tracking, the useful question is not which dashboard is prettiest. It is which number tells you what to do next.
What we measured, and how
On 26 July 2026 we ran 23 prompts through two surfaces, in the US market, in English. The prompt set was built the way a buyer actually searches: four branded prompts about our own company, and 19 non-branded prompts covering the questions our category gets asked — pricing, vendor selection, industry-specific automation, build-versus-buy.
Two signals, both machine-readable:
- Google AI Overview. For each query: does an AI Overview trigger at all, is our domain among its cited references, and which domains are cited instead. We also captured our classic organic position for the same query, which turned out to matter more than we expected.
- Embeddings retrieval. Whether our domain surfaces in the vector index layer that many AI applications query before generating an answer.
What we deliberately did not claim to measure: the answer text inside ChatGPT, Claude, or Perplexity. No API exposes what those systems said about your brand in a real user session. Any tool claiming to track this is running its own prompts through the API and inferring, which is a reasonable proxy but is not the same thing. We think vendors should say so plainly, so we are saying it about our own work first.
Finding one: AI Overviews are not an emerging surface, they are the surface
21 of 23 prompts (91%) returned a Google AI Overview. Not a handful. Not the informational ones only. Nearly every commercial question our buyers ask now gets answered above the organic results, by a system that cites a handful of sources and moves on.
That reframes what a ranking is worth. Position 4 on a query where an AI Overview occupies the first screen is not position 4 in any meaningful sense. If you are still reporting average position without segmenting by whether an AI Overview fires, your report is describing a page layout that most of your buyers are no longer seeing first.
Finding two: we scored zero, and the zero was the useful part
Here is the uncomfortable table. Of our four branded prompts, all four triggered an AI Overview and three cited us. Of our 19 non-branded prompts, 17 triggered an AI Overview and zero cited us.
A mention-counting tool would have reported "3 citations, 13% visibility" and drawn a line on a chart. That framing would have been actively misleading. The three citations were on queries containing our own company name — queries where a buyer already knows who we are. On every question where someone was actually shopping, we did not exist.
Branded visibility and non-branded visibility are different metrics measuring different business realities, and averaging them together produces a number that flatters you. If your AI visibility report shows a single blended percentage, split it. The branded number tells you whether your entity is understood. The non-branded number tells you whether you are in the consideration set. Only the second one is a growth metric.
Finding three: citation is mostly downstream of ranking
This is the finding that changed how we plan. We took six non-branded queries where an AI Overview fired, captured every domain it cited, and checked whether those same domains ranked in the classic organic results for the identical query.
Across 31 citations:
- 74% also ranked in the classic organic top 20
- 48% ranked in the top 10
- Only 25% were cited without ranking anywhere in the first two pages
The sample is small and we would not present it as a universal law. But the direction is clear enough to act on: for most queries, Google is drawing its AI Overview citations from a pool it has already decided to rank. There is no separate side door. The bulk of "AI visibility optimization" is ordinary ranking work, and a tracker that reports citations without reporting your organic position for the same query has hidden the causal variable.
That 25% residual is real and worth chasing — it is where format, structure, and being the definitive source on a narrow question earn a citation without a ranking. But it is the exception, not the strategy. Anyone selling you an AEO methodology that bypasses ranking should be asked to show the overlap data.
Finding four: one query type ignored websites entirely
One prompt in our set asked for the top AI automation agencies in a specific city. The AI Overview cited exactly four domains, and every one of them was a directory: Clutch, DesignRush, GoodFirms, and The Manifest. No agency websites at all. Not the market leaders, not the best-optimized pages. Directories.
For location-shaped queries, Google appears to treat curated third-party lists as the authoritative source and does not read vendor sites for the answer. If your visibility strategy for local intent is on-page content, you are optimizing an asset the system is not consulting. The lever is a complete, claimed directory profile.
We found this because we measured. We would not have found it by reading a best-practices article, and a mention-count dashboard would have shown a zero with no explanation attached.
Finding five: what actually gets cited
Across the 21 AI Overviews, 149 citations went to 116 unique domains. The concentration is lower than we expected — this is not a cartel of five big publishers.
User-generated and video content accounted for 27 of 149 citations, about 18%. YouTube appeared in 11 of the 21 answers, Reddit in 8, LinkedIn in 4. That is a meaningful minority but not the majority, which surprised us: the common claim that AI Overviews are dominated by Reddit did not hold in our category.
The remaining 82% was ordinary vendor and agency content. Which is the encouraging half of the result. The citation pool is not closed. It is populated by companies that rank, and ranking is a solvable problem.
What we would actually track
Based on running this, here is what belongs on an AI visibility report and what does not.
Track these:
- AI Overview trigger rate across your prompt set. This tells you how much of your category has moved above the fold. Ours is 91%; if yours is 20%, your priorities are completely different from ours.
- Branded and non-branded citation rates, reported separately. Never blended.
- Citation-to-ranking overlap. For every query where you are not cited, your own organic position for that query. This single column converts "we are invisible" into either "we need to rank" or "we rank and still are not cited," which are entirely different problems with entirely different fixes.
- The competitor citation map. Which domains are being cited instead of you, and how often. This is your real competitive set in AI answers, which frequently differs from your assumed competitive set.
Be sceptical of these:
- A single blended "AI visibility score." It compresses away every decision-relevant distinction.
- Claims to read actual ChatGPT or Claude conversation answers. No public API exposes them.
- Sentiment analysis on AI mentions, before you have any mentions to analyse. Sequence matters.
What this cost
The full 23-prompt run costs roughly ten cents in SERP API credit, plus the retrieval queries. It is a Python script that writes dated JSON, so month-over-month comparison is a diff rather than a screenshot.
We are not arguing nobody should buy a tool. Commercial trackers offer prompt libraries, scheduling, and reporting that a script does not. But you should know that the underlying measurement is cheap, and price the subscription against convenience rather than against access to otherwise-unavailable data.
How this connects to the work
We publish this partly because it is the same discipline we apply to client systems. In our AI consulting engagements the first deliverable is usually a measurement baseline, because a workflow nobody measured cannot be improved and an AI opportunity nobody sized cannot be prioritized. The audit that tells you where AI should not be applied is as valuable as the one that tells you where it should.
The same logic applies here. We ran this study on ourselves before recommending the approach to anyone, found that we score zero on every commercial query in our own category, and published it. If you are weighing whether you need outside help with AI strategy at all, our guide to what an AI strategy consultant does and when you need one covers the decision honestly, including when the answer is that you do not.
Frequently asked questions
What is AI visibility tracking?
Measuring whether AI answer engines — Google AI Overviews, ChatGPT, Perplexity, Claude — mention and cite your brand when people ask questions in your category. Useful tracking captures three things: whether an AI answer appears at all, whether you are cited, and which domains are cited instead of you. Tracking that reports only a mention count leaves out the diagnostic information.
Can any tool actually see what ChatGPT says about my brand?
Not in real user sessions. No public API exposes what a given assistant told a given user. Tools that report this are running their own prompts through the API and treating the results as a sample. That is a legitimate proxy and often a useful one, but it is a sample of what the model tends to say, not a record of what users were told. Be wary of any vendor that blurs the distinction.
Is AI visibility a separate discipline from SEO?
Mostly not, on our data. Across the citations we examined, 74% went to domains already ranking in the classic organic top 20 for the same query, and 48% to the top 10. That means the majority of AI visibility work is ranking work. A real but minority share of citations goes to pages that do not rank, which is where answer-shaped formatting and topical depth earn their keep.
How often should we re-measure?
Monthly is enough for most businesses, run on the same day each month with an unchanged prompt set so the comparison is meaningful. Changing prompts between runs is the most common way these reports become uninterpretable. Re-measure sooner only after a significant content or technical change you want to attribute.
We are cited on branded queries but not commercial ones. What does that mean?
That your entity is understood but you are not in the consideration set — exactly the pattern we found in our own data. It means the fix is not entity or schema work, which is already succeeding. It is competing for the commercial queries themselves: ranking for them, earning references from sources the answer engines trust, and being present wherever those queries are answered by third-party lists rather than vendor sites.
What is the cheapest way to start?
Write 20 prompts your buyers would actually type, split branded and non-branded, and check them manually once. It takes an afternoon and will tell you your trigger rate and roughly where you stand. Automate it only once you know the measurement is telling you something you will act on.
Measure before you optimize
We went into this expecting to learn which AI platforms mattered most. We came out with something more useful and less comfortable: a zero on every commercial query, a clear causal explanation for the zero, and one query type where the entire answer came from directories we had not claimed.
None of that came from a dashboard. It came from running the measurement, keeping the raw output, and being willing to publish the bad number. If you want to talk through what this would look like for your category, book a strategy call and bring the questions your buyers actually ask.