Sagentel
Home › Knowledge Base › How to Measure AI Visibility: The Metrics That Actually Matter
Measurement

How to Measure AI Visibility: The Metrics That Actually Matter

Rankings do not exist and clicks understate the value. The metrics worth tracking, how to collect them honestly, and how to report them.

9 min readBy Sagentel

Ask ten marketing teams how they measure their presence in AI search and you will get three answers. Some say they check ChatGPT manually every couple of weeks. Some point at a small and unstable line in their referral traffic report. Most admit they are not really measuring it at all.

That is understandable, because the obvious metrics do not work here. Rankings do not exist. Impressions are not reported. Clicks happen only sometimes, and the moments that matter most often produce no click whatsoever.

So you need a different measurement model. This article walks through the metrics worth tracking, how to collect them without fooling yourself, and how to turn them into a report that a CFO will accept.

Why the old dashboard breaks

Traditional search reporting rests on three pillars: position, impressions and clicks. Each one has a problem in AI search.

Position does not translate. An answer is a paragraph, not a list. There is no third place. You are either in the answer or you are not, and if you are in it, you might be the recommended option or the one mentioned as an afterthought.

Impressions are invisible. No assistant publishes how many people asked a given question. You cannot know the true volume behind any prompt, only your relative performance across the prompts you track.

Clicks understate value dramatically. When someone asks which agency they should hire and the answer names three, including you, that is a high value brand impression at the exact moment of consideration. Your analytics will record nothing.

The workable alternative is to measure your presence in a representative sample of conversations, the way a brand tracking study measures awareness rather than counting shop visits.

Metric 1: Presence rate

What it is: the percentage of your tracked prompts where your brand appears anywhere in the answer.

This is your headline number. If you track 60 prompts and appear in 21 of them, your presence rate is 35 percent.

Why it matters: it answers the most basic question anyone will ask you, which is whether you show up at all.

How to read it: the absolute number means little without context. Presence rate in a category with four competitors looks nothing like presence rate in a category with two hundred. What matters is the trend and the comparison to competitors.

The trap: measuring presence across your whole prompt set only. An overall figure of 35 percent can easily hide a situation where you appear in 80 percent of brand prompts, which is expected and unimpressive, and 6 percent of solution prompts, which is where deals are actually won. Always break presence down by cluster.

Metric 2: Share of voice

What it is: your presence rate compared against named competitors across the same prompt set.

Why it matters: this is the number that gets budget approved. Executives may not know what a good presence rate looks like. They immediately understand that a competitor appears in twice as many answers as you do.

How to calculate it: for each prompt, record every brand mentioned in the answer. Then compute each brand's share of total mentions across the set. If 60 prompts produce 180 brand mentions and you account for 27 of them, your share of voice is 15 percent.

How to read it: watch which competitors show up that you did not expect. AI answers regularly surface smaller, content heavy players that never rank well in traditional search, because they happen to have written the clearest comparison page in the category. Those competitors are your real benchmark, not the market leader.

Metric 3: Citation rate

What it is: how often your own domain is used as a linked or named source, as opposed to your brand simply being mentioned.

Why it matters: presence and citation are influenced by different work. Presence is driven mostly by what the whole internet says about you. Citation is driven by whether your own pages are retrievable and quotable. Citation rate is therefore the metric most directly under your content team's control, which makes it the best measure of whether your content programme is working.

What to track alongside it: which specific URLs get cited. This is genuinely useful data. You will usually find that citations concentrate in a small number of pages, and that these are rarely your most polished marketing pages. They tend to be documentation, comparison pages, glossary entries and data posts. That tells you exactly what to make more of.

Metric 4: Accuracy

What it is: the proportion of answers about your brand that describe you correctly.

Why it matters: presence without accuracy can be actively harmful. If an assistant confidently tells buyers you do not integrate with a platform you have supported for two years, high visibility is making things worse.

What to check: pricing, positioning, feature claims, target market, company facts like size and location, and any claim about limitations. Score each brand answer as accurate, partly accurate or wrong, and log the specific error.

What to do with errors: trace them. Wrong claims almost always come from an identifiable source, often an outdated review, an old comparison article, a stale press release or your own neglected page. Fixing the source is the only real remedy.

Metric 5: Sentiment and position within the answer

What it is: whether you are described positively, neutrally or negatively, and whether you appear as the primary recommendation or as a secondary option.

Why it matters: two brands can both have a presence rate of 40 percent while having completely different commercial outcomes, because one is consistently listed first with a positive framing and the other is mentioned last with a caveat about being expensive or complicated.

How to score it: keep it simple. Position as first, middle or last mention. Sentiment as positive, neutral or negative. Resist the urge to build a ten point scale, because the extra precision is not real and it makes the tracking too laborious to sustain.

Metric 6: Answer variance

What it is: how consistent the answer is when the same prompt is run repeatedly.

This one is usually ignored and it matters more than people expect. Assistants are probabilistic. The same prompt asked five times can produce five different sets of recommended vendors. If you check a prompt once and record a result, you have a data point with an error bar wide enough to swallow any change you were hoping to detect.

Run each tracked prompt at least three to five times, on different days, and record presence as a rate rather than a yes or no. A brand that appears in four of five runs is in a genuinely stronger position than one that appears in one of five, even though a single check might show both as present.

Variance is also a useful signal in its own right. Low variance means the model has a settled view of your category. High variance means the answer is unstable, which is an opportunity, because unstable answers move faster in response to new content.

Metric 7: Downstream signals

These are the numbers your finance team will eventually ask about. Treat them as directional rather than precise.

Referral traffic from assistant domains. Real but understated, since many assistants do not pass a clean referrer. Segment it in your analytics and watch the trend rather than the absolute number.

Branded search volume. If AI answers are introducing you to new buyers, some of them will search your name afterwards. A rising branded search line alongside a rising presence rate is a reasonable correlation to point at.

Self reported attribution. Add "an AI assistant" as an option to your "how did you hear about us" field. It is the crudest method available and it is also, in practice, one of the most informative.

Conversion quality. Traffic arriving from AI assistants tends to be further along in the buying process, because the assistant has already done the filtering. Compare conversion rates by source rather than judging the channel on volume.

Building a report people will read

Keep it to one page, updated monthly, with five elements.

A single presence rate figure with the month over month change.

A share of voice chart with you and three named competitors.

A cluster breakdown showing presence by buying stage or product line, which is where the actionable insight lives.

An accuracy note listing any new factual errors found and what is being done about them.

A short list of the pages currently earning citations, so the content team can see what is working.

Everything else belongs in an appendix that nobody will open, and that is fine.

Setting targets without inventing numbers

Do not commit to a presence rate target in your first quarter. You have no baseline and no sense of the natural variance, so any number you promise is a guess that will be held against you.

Spend the first quarter establishing the baseline and measuring variance. From the second quarter, set targets as relative improvements against your own baseline and against a named competitor. Something like "close the share of voice gap with Competitor B from 14 points to 7" is defensible. "Reach 60 percent presence" usually is not, because you do not control the denominator.

Also set a target on accuracy, because it is the one metric where a specific goal is entirely reasonable. Zero known factual errors in brand answers is an achievable standard.

A monthly rhythm that is actually sustainable

Week one: run the full tracked prompt set, multiple runs per prompt. Record presence, competitors mentioned, sources cited, sentiment and any factual errors.

Week two: analyse by cluster. Identify the weakest cluster and the pages that would most plausibly fix it.

Weeks two to four: do the work. Publish, update, fix sources, chase reviews, pitch data.

Week four: report, using the same one page format every month so trends are visible at a glance.

The discipline that matters most is keeping the prompt set stable. If you change the questions every month you will produce numbers that move constantly and mean nothing. Lock the core set, rotate only the experimental set, and let the trend line do its work.

The point of all this

Measurement here is not about proving a channel deserves budget, although it does help with that. It is about being able to see a market you currently cannot see.

Right now, buyers in your category are asking assistants which company to work with, and those conversations are producing shortlists you never observe. Presence rate, share of voice and accuracy are the closest thing available to a window into that room. Even a rough view of it beats the alternative, which is finding out you were left off the shortlist when the deal is already gone.

Track your brand in AI answersSagentel replays the questions your buyers ask ChatGPT, Gemini, Perplexity and Claude, and shows you the answers, the sources and what to do next.
Start free
Keep reading