1001 SEO Media
All posts
By 1001 SEO MediaGEOshare of voiceAI visibilitymeasurement

AI Search Share of Voice: Measure Brand Mentions in Generative AI Answers for Visibility

How we define AI share of voice in a client scope: the prompt set, the denominator, mention versus citation, position, sentiment, and how often we sample answers.

Most clients want one number for AI visibility, and share of voice is the one they ask for by name. We are happy to give it. We just write the definition into the scope of work before we report it, because a share of voice figure without its definition can be pushed up or down by anyone who controls the inputs, including us.

This post is that definition: how we calculate share of voice in generative AI answers, which choices change the number, and how we run the measurement on client retainers.

The formula we put in the scope

Our working definition is simple. Across a fixed set of prompts, checked in one engine over one period, share of voice is the client's brand mentions divided by the total mentions of every brand in a named comparison set.

A made-up example shows the arithmetic. Say we check 40 prompts in ChatGPT. The answers name the client 18 times and the five listed competitors 72 times between them. The client's share is 18 out of 90, or 20 percent.

We report a second number next to it, which we call presence rate: the share of prompts where the brand appears at all. In the same example the client might appear in 15 of the 40 prompts, a presence rate of 37.5 percent. The two numbers move for different reasons. Presence rate rises when you break into answers you were missing from. Share of voice rises when you take room from competitors inside answers you were already in. A client who only sees one of them will misread the other's movement.

The comparison set decides the number

Share of voice only means something relative to the brands in the denominator. Add a small competitor that rarely gets named and the client's share goes up without a single answer changing. Remove the category leader and the same thing happens, more dramatically.

So we fix the comparison set at kickoff, usually the brands the client loses deals to plus whoever the engines name most often in the first baseline run. Any later change gets logged with a date, and we restate the prior period on the new set so the trend line stays honest. If a client asks us to drop a competitor because "they're not really a competitor," we will, but the report says so on the page where the number appears.

Fix the prompt set before you measure anything

The prompt set is the other half of the denominator. We build it from how buyers actually ask: category questions, comparison questions, problem questions, and a few prompts that name a competitor. Branded navigational prompts ("what is [client]") are excluded from share of voice. They inflate it, and the client does not need a tool to know ChatGPT can describe them when asked by name.

Brand spelling matters more than people expect. If the engines refer to the client by an abbreviation, a former name, or a product name, those need to count as mentions. We set up aliases before the first run, which is covered in our note on aliases and the visibility heatmap. Missing an alias is the most common reason a share of voice number looks wrong in month one.

Once the set is live, we change it as little as possible. Swapping ten prompts mid-quarter makes the quarter incomparable with the last one, and no chart annotation fully fixes that.

Four columns that stay separate

A mention is not a citation. The model can name a brand without linking to it, and it can cite a brand's page while naming a competitor as the recommendation. We count mentions for share of voice and track citations alongside, per domain and per page, because the work that improves each is different. Citation gaps usually lead to content and offsite work. Mention gaps more often lead to positioning and third-party coverage.

Position is the third column. Being named first in a list and being named fifth both count as one mention in the formula above. We do not fold position into a weighted score in our own reporting. We show average position next to share of voice and let the reader see both. Weighted composites are fine inside a platform, as long as everyone reading the slide knows the weighting.

Sentiment is the fourth. A mention with a warning attached ("cheaper, but support is slow") is still a mention. We count it, then flag it in the sentiment column so it does not get celebrated as a win.

Per engine, per market, per persona

A blended share of voice across engines hides more than it shows. ChatGPT has 820M+ weekly active users; Perplexity has 22M+ monthly. If you average the two engines equally, Perplexity counts far more than its audience. If you weight by users, Perplexity almost disappears. Neither choice is wrong, but it is a choice, and the client should make it knowingly. Our default is to report each engine on its own line and give one blended figure for orientation, with the weighting written underneath.

The same applies to markets and buyer types. A brand can lead share of voice in one country and trail in another, and a blended number shows neither. We split by country (and by city for local businesses) and by persona when the buyer types ask different questions.

Sampling cadence and noise

AI answers vary between runs. The same prompt can name four brands on Monday and three on Tuesday. One check is an anecdote, so we only report share of voice over repeated checks, and we read trends over weeks, not days.

When every tracked prompt shifts on the same day, we look for a model update before we look for a cause in our own work. When one topic shifts after we shipped a page for it, that is worth investigating, though it still is not proof.

How we run it in Promptwatch

The platform we run client programs on is Promptwatch, and share of voice is one of the first views we open. It computes share of voice and competitive benchmarking across the fixed prompt set, with personas and country, state, or city tracking for the splits above. Prompt trends show how a prompt's visibility moved over time and what changed between checks, which is what we need when a client asks "why did we drop on that prompt." Citation analytics sit next to it, so the mention column and the citation column come from the same set of monitored answers.

Two more pieces help with the prompt set itself. Prompt search volumes and difficulty scores tell us which prompts are worth a slot, and query fan-outs show the sub-searches an engine runs behind a prompt, which explains why a citation can land on a page that does not match the typed question.

Sizing is straightforward. Essential at $95 a month holds 50 prompts for one project. Professional at $245 holds 150 prompts across 2 projects, which is where most of our single-brand share of voice programs sit once a second market is added.

What we refuse to report

We will not report a single blended share of voice without the scope block beside it: engines, markets, prompt count, comparison set, period. We will not present a share of voice rise that came from editing the prompt set or the competitor list as progress. And we will not translate share of voice into revenue. It measures how often you are named relative to rivals in a defined sample. That is useful on its own, and it does not need to be dressed up as something else.