Measuring AI search visibility
You cannot manage AI visibility with the metrics SEO left you, because ranking position does not exist in an AI answer. Five metrics track it instead. Mention rate: how often an engine names you for the prompts that matter. Citation rate: how often it links you as a source, not just a name. AI share of voice: your slice of those answers against your competitors. Position in the answer: where you sit in the reading order the model composed, since the first source named carries more weight than the fifth. AI referral traffic: the visits that actually arrive from the answer. The hard part is that each must be measured per engine and per prompt set, because presence on ChatGPT does not imply presence on Perplexity, Claude or Gemini, and a brand strong on one can be absent on another. Rolled together, the five become a single GEO Score from 0 to 100 a non-technical owner can read at a glance, and Surface acts on the gaps behind that number across all five engines so the score turns into published work. This guide defines each metric, explains the per-engine and per-prompt traps, and shows how to turn the numbers into a weekly loop.
If you came from SEO, your instinct is to open a rank tracker and check your position. In AI search that instinct fails on contact, because there is no position to check. The engine read the sources for you and returned one answer. The only questions left are whether you are in it, how often, how you are framed, and where you sit.
This guide defines the metrics that answer those questions, names the traps that make them easy to measure wrong, and shows how to turn them into a loop. No borrowed numbers: where a figure would matter, we point to our own ongoing measurement rather than cite one we cannot stand behind.
Why the old metric does not transfer
Ranking position measured a list. A person scanned ten links, and being third instead of eighth changed whether they clicked. The metric existed because the surface was a list and attention dropped down it.
An AI assistant removes the list. It composes one answer and names a few sources inside it. There is no eighth place to climb out of, so position on a results page stops meaning anything. The unit of visibility moved from a rank on a page to a presence in a paragraph, and the metrics have to move with it.
That is not a tooling gap you can patch with a better rank tracker. It is a different surface, and it needs metrics built for the answer, not for the list the answer replaced.
The five metrics that track AI visibility
Five numbers describe whether an engine names, trusts and recommends you. Read in order, they go from "are you there at all" to "did it produce anything".
| Metric | What it answers | Why it matters |
|---|---|---|
| Mention rate | How often an engine names you for the prompts that matter | The binary floor: you cannot be recommended if you are not named |
| Citation rate | How often it links you as the source of a claim | Stronger and more durable than a bare mention |
| AI share of voice | Your slice of those answers versus competitors | Visibility is relative; a rising tide can still leave you behind |
| Position in the answer | Where you sit in the model's reading order | The first source named carries more weight than the fifth |
| AI referral traffic | The visits that actually arrive from the answer | The downstream proof that presence turned into action |
Mention rate is the floor. For a defined set of prompts your buyers actually ask, how often does the engine say your name? If it is zero, nothing below matters yet.
Citation rate is the next rung. A mention is your name in prose; a citation links you as the source behind a claim. Citations are stronger, harder to fake, and more durable, because they reflect the model trusting you as a reference rather than recalling you in passing.
AI share of voice makes it relative. Visibility is a competition for a finite answer, so the right comparison is not "are we mentioned" but "what fraction of the answers in our category name us versus a competitor". You can improve in absolute terms and still lose share if a rival improves faster.
Position in the answer captures framing. Being named first, as the recommended option, is worth more than being listed fifth as an also-ran. The reading order the model composed is the closest thing the answer layer has to a rank, and it is qualitative as much as ordinal.
AI referral traffic is the proof. The visits that arrive from an assistant's answer are the downstream signal that presence became action. It is the smallest and noisiest of the five today, because attribution from inside an answer is still maturing, so weigh it as confirmation rather than as your primary dial.
The two traps: per engine, per prompt
The metrics are simple. Measuring them honestly is where most efforts go wrong, in two specific ways.
First, measure per engine. Presence on ChatGPT tells you nothing about Perplexity, Claude or Gemini. Each pulls from different sources and composes differently, so a brand strong on one can be absent on another. An average across engines hides exactly the gap you need to see. Keep the score per engine, then aggregate, never the reverse.
Second, measure per prompt set. "Are we visible in AI" is not a measurable question. "How often is our brand named when someone asks for the best tool in our category, across these forty prompts, on each engine" is. The prompts have to be stable so movement is comparable week to week, and representative so you are tracking the buying questions that matter rather than vanity queries you happen to win.
Get either wrong and the number lies. A flattering average across engines, or a prompt set quietly tuned to your strengths, produces a score that goes up while your real visibility does not.
Rolling it into one score
Five metrics across many engines and prompts is the right level of detail for the work and the wrong level for the owner who needs to know, in one glance, whether things are improving. So they roll into a single GEO Score from 0 to 100, and Surface reads all five across ChatGPT, Perplexity, Claude, Gemini and Google AI Overviews and acts on the gaps behind the number.
The score is the headline. Its value is that you can drill from it: a drop is not "we went from 62 to 54", it is "Perplexity stopped citing us on three buying prompts because a competitor published a comparison". The number gets attention; the breakdown drives the decision. Kept that way, the score is a steering instrument, not a vanity metric.
Turn the numbers into a loop
A measurement you take once is a report. A measurement you take on a cadence is a system. AI answers shift as engines recrawl, as competitors publish, and as the models update, so a single snapshot ages within weeks.
So run it as a loop: measure the prompt set per engine, see what moved, change one thing, and measure again. A weekly cadence on a stable prompt set is enough to catch real movement and react when a competitor takes your slot, without chasing daily noise. That continuous loop, run across every engine that matters, is the whole discipline.
Surface Agent runs that loop for you. It watches your visibility across ChatGPT, Perplexity, Claude, Gemini and Google AI Overviews around the clock, along with the Reddit threads where your buyers ask for recommendations. It drafts the content and comparison pages that move the metrics, and drafts the Reddit replies in your voice for you to post from your own account. You review and approve, and the score climbs.
Frequently asked questions
01Why can't I just use my SEO rankings to measure AI visibility?
Because there is no ranking in an AI answer. The engine returns one composed response and names a short list of sources, so the question is not what position you hold but whether you are named at all, how often, how you are framed, and where you sit in the reading order. SEO rank tools measure a results page that the answer layer replaced; they cannot see inside the answer.
02What is the single most important AI visibility metric?
There is no single one, which is the point, but if forced to pick, citation rate per engine on the prompts that matter most. A mention puts your name in the answer; a citation links you as the source behind a claim, which is stronger and more durable. Measure it per engine, because being cited on ChatGPT tells you nothing about Perplexity, Claude or Gemini.
03How often should I measure?
Treat it as a continuous loop, not a one-time report. Answers shift as engines recrawl, as competitors publish, and as the models update, so a single snapshot ages fast. A weekly cadence on a stable prompt set is enough to see real movement and react when a competitor takes your slot, without chasing day-to-day noise.
04What is a GEO Score?
It is a single number from 0 to 100 that rolls the underlying metrics (mention rate, citation rate, share of voice, position and referral signal) into one figure a non-technical owner can track at a glance. The score is the headline; the value is in drilling from it into which engine and which prompt moved, so it stays a decision tool rather than a vanity number.