Measuring AI visibility without fooling yourself

6 min read
  • AI visibility
  • Measurement
  • GEO

The first thing most teams do is ask ChatGPT whether it recommends them, get a pleasing answer, and conclude they are visible. Ask again an hour later and the answer may name three competitors instead. Neither run was wrong; both were samples of one.

This is the central difficulty in measuring AI visibility, and almost every bad AI visibility metric comes from ignoring it. Answer engines sample when they generate, their retrieval sets change as indexes update, and many of them vary by region or session. A number built on one query per question is mostly measuring that variance.

Start with a fixed question set

Your unit of measurement is not a keyword, it is a question a buyer would actually type. “project management software” is a keyword. “What’s the best project management tool for a small agency that bills hourly?” is what someone asks an answer engine, and the second one is what you can track.

Write thirty to fifty of them, spanning the decision: category definition, comparisons, alternatives-to queries, objection-shaped questions like “is X worth it”. Then freeze the set. The moment you add and remove questions between runs, your trend line stops meaning anything, because you have changed the instrument and the reading at the same time.

The discipline that mattersKeep the question set, the engines, and the phrasing constant. Change one thing at a time, and never the instrument and the site in the same week — otherwise you cannot attribute the result to either.

Three metrics that hold up

Citation share

The share of your tracked questions where your domain is cited. It is the headline number, and it is only meaningful against a fixed question set. Report it as a fraction with the denominator visible — “6 of 40” resists over-reading in a way that “15%” does not.

Competitive displacement

For every question where you are absent, which domains are cited instead. This is the most actionable thing you can collect, because it is a list of pages an engine currently prefers over yours for a question you care about. Read the winning pages and the pattern is usually visible within a dozen of them.

Answer position

Whether you are the answer’s primary source or a supporting citation near the end. Being cited fourth in a list of five is a materially different outcome from being the source the answer is built on, and collapsing both into “cited” hides real movement.

Numbers to distrust

  • Anything from a single run. One query is a sample. Repeat each question several times per measurement period and record the rate, not the last answer.
  • Sentiment scores on generated text. An engine’s phrasing about you varies run to run far more than its citations do. The signal-to-noise ratio is poor enough that the number moves mostly at random.
  • “Visibility scores” with no stated denominator. If you cannot see the question set and the engine list, the number is not comparable to anything, including its own previous value.
  • Self-reported model knowledge. Asking a model whether it knows your brand tests its willingness to agree, not its retrieval behaviour. Ask the buyer’s question and read the citations instead.

Cadence and expectations

Weekly is usually the right rhythm. Daily measurement on a non-deterministic system mostly produces a jagged chart you learn to ignore, and monthly is too slow to connect a change to its cause.

On timing: after you publish or restructure a page, the engines have to recrawl it, and the ones that lean on a search index have to wait for that index to update. Weeks, not days. If your citation share moves the morning after a deploy, be suspicious — you are more likely looking at variance than at the effect of your change.

One asymmetry worth planning around: losing visibility shows up faster than gaining it. A competitor publishing a definitive page can displace you within a crawl cycle, while earning the citation back takes the full recrawl-and-corroborate path. Measurement is how you notice the first case in time to respond.

A workable first week

  1. Write 30 buyer questions and store them somewhere they will not be quietly edited.
  2. Run each three times across the engines that matter for your category, recording every cited domain.
  3. Compute citation share and the displacement list. The displacement list is your content backlog, already prioritised by an engine.
  4. Change one thing — usually restructuring the two or three pages closest to your weakest questions.
  5. Re-run the identical set two to three weeks later and compare. Same questions, same engines, same phrasing.

Done by hand this is a few hours a week, which is exactly why it tends to get abandoned after the second run — and a trend you stopped collecting is worth nothing. If you would rather not maintain the spreadsheet, our free AI visibility audit runs a question set against the major engines and reports the displacement list, and the platform guide explains how the tracking runs on a schedule.

Common questions

What is citation share?
Citation share is the proportion of your tracked questions where an answer engine cites your domain. If you track 40 questions and your domain appears in 6 of the answers, your citation share is 15 per cent. It is only comparable over time if the question set stays fixed.
Why do AI engines give different answers to the same question?
Generation is sampled rather than deterministic, retrieval sets shift as indexes update, and most engines personalise or vary by region. The practical consequence is that a single query result is a sample, not a measurement, and conclusions drawn from one run are unreliable.
How many questions should I track?
Enough that one question changing its answer does not move your headline number much. Below roughly 20 questions, normal run-to-run variance is larger than most real changes you are trying to detect. Thirty to fifty buyer-intent questions is a workable range for a single category.
Should I track brand mentions or only linked citations?
Track both, separately. A linked citation is the stronger outcome and the one that can send traffic. An unlinked mention still shapes what the buyer believes, and it often appears before citations do, which makes it a useful early indicator.

Read next