After this lesson you should be able to
- List what can and cannot be measured reliably about generated answers
- Build a repeatable prompt panel and record results defensibly
- Combine proxy metrics that are stable with direct metrics that are not
- Challenge vendor claims about AI visibility scores
Start with what is broken
Search measurement took two decades to get to rank tracking, impression data and a defensible attribution model, and it is still argued about. Measurement of generative answers is roughly where search measurement was in 1999, and I would rather say that plainly than sell you a dashboard.
Four problems, none of them solved:
- Non-determinism. The same prompt can produce different answers, with different sources, on consecutive runs. A single observation is an anecdote.
- No sampling frame. Rank tracking works because keyword data gives you a rough distribution of what people search. Nobody publishes an equivalent distribution of what people prompt. You are sampling an unknown population, which means your share of voice has no denominator.
- Personalization and context. Answers can vary by account, history, location, device and conversation state. What you see is not what a customer sees.
- The systems change under you. Models are updated, retrieval is retuned, interfaces are redesigned. A trend line across a version change is two different measurements plotted on one axis.
Add one more from Lesson 9.4: a citation is a visible proxy, not a full accounting of what influenced an answer. So even a clean count of citations is measuring the shadow rather than the object.
What can be measured usefully
Given all that, here is what I think is genuinely worth doing. None of it is precise. All of it is directional, and directional is not nothing.
A fixed prompt panel. Write 30 to 60 prompts that represent how a real buyer would ask about your category — problem-first questions, comparison questions, “best X for Y” questions, and questions where you would expect to be named. Freeze the list. Run it on a schedule, in a clean session, against the same set of systems. Record the full verbatim answer, the sources shown, the date and the system version. The value is entirely in doing it the same way every time.
Presence and share of voice within that panel. Across your frozen prompts, how often are you named at all, how often are you cited as a source, and which competitors appear instead? Report it as “across our 40 tracked prompts, we were named in 12” — never as a percentage of some imagined universe of prompts.
Accuracy and sentiment of what is said. Underrated and often the most valuable output. Is the description of your product correct? Is it citing a page you retired? Is it repeating a competitor's framing of you? Factual errors about your business are a concrete, fixable problem that shows up nowhere else in your reporting.
Mention volume and quality. Media monitoring on your brand and key people, scored for topical fit and source independence, as Lesson 9.2 describes. This is stable, measurable and available now.
The proxies that are more reliable than the direct measures
Slightly counterintuitively, the sturdiest evidence that this work is landing sits in metrics you already have.
- Branded search volume and direct traffic. If more people are being told about you, more of them look you up. This is slow, noisy and confounded by everything else you do — but it is real data on real people, not a scraped answer panel.
- Referral traffic from assistant and AI-search domains. Your analytics will show some of it by referrer. Treat it as an undercount: not every surface passes a referrer, some traffic arrives as direct, and the labelling changes without notice.
- Assisted conversions and self-reported attribution. A “how did you hear about us” field on the form is crude, unfashionable, and often the only place this shows up at all.
- Placement and mention counts from your own PR reporting, which you control and which do not change definition overnight.
I would rather brief a board on branded search trend plus a documented prompt panel than on a vendor's composite visibility score. One of those I can explain the derivation of.
The attribution problem underneath all of this
There is a harder problem sitting beneath the metrics, and it is worth naming because it explains why the measurement gap is not going to close quickly.
When a generated answer describes your product accurately and the reader decides you are worth considering, the influence happened in a place you have no instrumentation on. There may be no click. If a click follows later, it will often arrive as a branded search or a direct visit, and your analytics will credit whichever channel it can see. The mechanism that produced the demand is invisible; the channel that harvested it takes the credit.
This is not a new phenomenon — it is exactly what happened with featured snippets, with review sites, and with word of mouth long before either. It is only newly annoying because it is growing. The practical response is the unglamorous one: watch the demand-side indicators rather than trying to instrument the invisible step. If branded search is rising, direct traffic is rising, and sales conversations start with an accurate description of what you do that nobody on your team supplied, something is working upstream of your tracking.
Do not let anyone sell you a model that claims to resolve this. Nobody has resolved it for word of mouth in a century of trying.
How to challenge a vendor claim
A market of monitoring tools has appeared, and some of it is good work. Some of it is a number invented to be sold. Four questions separate them, and you should ask all four before you buy.
- Where does the data come from? API access, or an automated browser session? These produce different results, and a scripted session is not what your customers see.
- How many runs per prompt, and do you report variance? If a tool reports a single figure per prompt per day with no spread, it is hiding non-determinism rather than measuring it.
- What is the denominator? Any “share of voice” percentage needs a defined prompt set. If the prompt set is theirs and undisclosed, the percentage is unauditable.
- What happens on a model update? Ask how they handle discontinuities. If historic data is silently re-based, your trend line is fiction.
A tool that answers all four honestly is worth paying for. A tool that gives you a single score out of 100 with no methodology is a confidence trick with a nice interface.
What I would actually build this quarter
Something deliberately modest, because modest and repeatable beats sophisticated and abandoned.
- Write 40 prompts. Freeze them. Store them in the repository with your other reporting assets.
- Run them monthly, three times each, in clean sessions, across the two or three systems your audience plausibly uses. Log verbatim answers, sources, date, system.
- Record three numbers: prompts where you are named, prompts where you are cited, prompts where a competitor is named and you are not.
- Log every factual error about your business and treat each one as a content ticket.
- Chart branded search and referral traffic alongside it, and never present the panel results without them.
- Write the caveats into the report itself, in the report, every month. Not in a footnote.
That is not a measurement system. It is a documented observation practice, which is the honest thing to have while the field settles. If someone asks me in two years whether these numbers were rigorous, I want the answer to be that we said so at the time.
Questions
Can I track my ranking in AI answers the way I track keyword rankings?
No, and the analogy is misleading. Rank tracking works because results are ordered, largely stable, and keyword data gives a rough demand distribution. Generated answers are non-deterministic, unordered, personalized and prompted in ways nobody has a distribution for. You can track presence across a fixed prompt panel, which is useful and much weaker than rank tracking.
Is AI referral traffic showing up in my analytics?
Partly. Some surfaces pass a recognizable referrer and some do not, chat clients often arrive as direct traffic, and the labelling changes without announcement. Segment what you can identify and treat it as a floor rather than a total. Do not build a business case on the precision of a figure you know is incomplete.
How many prompts should a monitoring panel contain?
Enough to cover the real question types and few enough that you will actually run it every month — 30 to 60 works for most businesses. Cover problem-first questions, comparisons, category questions and the ones where you would expect to be named. Freeze the list, because changing it breaks comparability, which is the only thing the panel really gives you.
Should I buy an AI visibility tool?
Possibly, if it answers four questions honestly: where the data comes from, how many runs per prompt and with what variance, what the denominator of any percentage is, and what happens to history when a model updates. A single opaque score out of 100 is not measurement. A documented, repeatable prompt panel — even a manual one — is.