Part 1 · Foundations · Updated September 13, 2026 · 6 min read
How to Measure GEO: Mentions, Citations and Stability
Everything before this chapter argued that AI visibility is real and improvable. This chapter is about knowing, rather than hoping, that yours improved. It is where most GEO programs quietly fail, because the channel publishes no data about itself: no query volumes, no impressions, no position reports. Whatever you know, you measured yourself.
The five numbers that describe AI visibility
Mention rate. Of many runs of a buying question, in what share of answers does your brand appear at all? This is the fundamental unit: a rate across runs, never a yes or no from one run.
Share of voice. When the engine names brands for your topic, what fraction of those namings are you, against each competitor? Visibility is relative; being mentioned in 40% of answers means one thing in a field of two and another in a field of nine.
Citation rate. How often is your own site the source the answer links to? You can be talked about through other people’s pages. Fine start; fragile position, because you do not control what those pages say next.
Sentiment and accuracy. When you do appear, what does the answer actually say? Recommended, listed with caveats, or described wrongly? An engine confidently misstating your pricing is visibility working against you.
Stability. How much do all of the above move between runs, phrasings and weeks, when nothing about you changed? This is the metric almost nobody reports, and it is what makes the other four readable. A 60% mention rate that swings twenty points on rerun is a different fact than a steady one.
Why one reading is not a measurement
Chapter 3 established that an engine holds a distribution of answers for every question. Measurement, then, is sampling: run the same prompt set many times, across phrasings and personas, and count.
Sampling has a property that surprises people used to rank trackers: every number comes with uncertainty, and honest reporting says how much. If you ran a prompt five times and appeared in three answers, your mention rate is not “60%”. It is “roughly somewhere between weak and strong”, because five runs cannot pin it tighter. More runs narrow the range. The range, not the point, is the measurement.
This has a blunt consequence for reading dashboards, yours or any vendor’s: a score that moved is not news until the movement is larger than the score’s normal wobble. Teams that skip this step spend their quarters reacting to noise, celebrating recoveries from dips that never happened, and losing faith in the channel when the “gains” evaporate.
Averages hide; breakdowns inform
A single overall score compresses away everything actionable. The same brand, honestly measured, might be:
- strong in comparison questions, absent in awareness questions,
- recommended to startups, never mentioned to enterprise buyers,
- visible on ChatGPT, invisible on Gemini,
- solid in English, missing in every other language its buyers use.
Each of those is a to-do list item. The average of them is a mood. Insist on seeing every score broken down by topic, persona, engine and market, and treat any tool or report that only offers you one number with suspicion.
Measuring fixes like experiments
The point of measurement is to steer work. The loop looks like this:
- Baseline. Measure your prompt set properly before touching anything, so “before” is a distribution, not a screenshot.
- Change one thing deliberately. Rewrite the page, add the evidence, fix the wrong fact, publish the comparison. Chapter 4’s tiers say where to start.
- Re-measure the same way. Same prompts, same method, enough runs. Different measurement setups before and after tell you about the setups, not the fix.
- Believe the change only if it clears the noise. If the lift is within the normal wobble, the honest verdict is “not proven yet”, and that verdict protects you from building strategy on luck.
Run this loop and your GEO program compounds: you accumulate not just visibility but knowledge of which levers move it for you. Skip it and every month starts from opinion. Part 2 of this guide turns the loop into an operating rhythm.
Full disclosure of our interest here: this measure-with-uncertainty approach is what Relevant is built to do, sampling your buying questions across phrasings, personas and engines, and reporting every number with its confidence range. But the method is not proprietary. The mechanics in this chapter can be done with a spreadsheet and patience, and chapters 6 and 7 show exactly how, along with the one thing a spreadsheet cannot tell you: whether you were asking the right questions in the first place.
Questions people ask about this
How many runs are enough?
Enough that the range around your number is narrower than the decisions you want to make with it. As a floor, several runs per phrasing; the more the answers wobble, the more runs the topic needs. What is never enough is one.
What is a good AI share of voice?
There is no universal benchmark; the channel is too new and too category-dependent. The usable standard is your own trajectory against your own competitors, measured the same way every time.
Which metric should we report to leadership?
Share of voice on your money topics, with its trend and its uncertainty, plus sentiment on anything the engines get wrong about you. One honest chart beats five precise-looking ones.
Can we just check ChatGPT manually every Monday?
A manual weekly spot-check is one run of a handful of prompts: better than nothing, and far too thin to act on. If manual is the constraint, chapter 7's audit structure gets the most signal out of it.