Skip to content

Blog

How to measure AI visibility: a repeatable method

Mike Holp · Published · Updated · Reviewed · 6 min read

To measure AI visibility you ask AI answer engines the questions your buyers actually ask, repeat each question enough times to see past the noise, and record three things every run: whether you were named, who was named instead, and which sources were cited. Everything else is refinement.

The reason this needs a method at all is that AI answers are non-deterministic. Ask the same question twice and the names can change. A single prompt is an anecdote; a measurement is what you get when you sample.

Step 1: write buyer-intent prompts, not brand prompts

The most common mistake is asking "what do you know about [my company]?" That measures whether the model has heard of you, which is not the same as whether it recommends you.

Ask instead the way a buyer with a problem would ask, without naming any vendor:

  • "Who are the best [category] providers in [location]?"
  • "What should I use to [job the buyer wants done]?"
  • "Which [category] tool is best for [specific constraint]?"

If your name only appears when you put it in the prompt, your visibility for buyer-intent questions is zero, regardless of what a brand prompt returns.

Step 2: sample repeatedly

Run each prompt several times rather than once. Three to five runs is usually enough to tell a stable result from a coin flip. What you are looking for is a rate, not a fact: named in four of five runs is a meaningfully different position from named in one of five, and a single run cannot distinguish them.

Keep temperature and phrasing constant between runs so the only variable is the model's own non-determinism.

Step 3: cover more than one engine

ChatGPT, Claude, Perplexity, and Gemini draw on different sources and disagree regularly. Measuring one engine gives you a quarter of the picture with no indication of which quarter. If you must pick one, pick the one your buyers tell you they use — but treat the result as engine-specific rather than as "AI".

Step 4: record the same fields every time

For each run, record:

FieldWhy it matters
Engine and dateRetrieval behaviour changes; an undated result is not comparable to anything
Prompt, verbatimSmall rewordings change answers, so the prompt is part of the measurement
Named or notThe primary outcome
Competitors namedUsually the most actionable output — who the engine treats as the safe answer
Sources citedWhere a mention would actually change the answer
Quoted answer textEvidence, so a number can be re-checked later

The quoted text matters more than it looks. A score with no evidence behind it cannot be audited, argued with, or diffed against next month's run.

Step 5: turn runs into three numbers

Mention rate — the share of runs naming you, per engine, per prompt. Competitor share — which other businesses appear, and how often. Citation set — the domains cited across all runs, ranked by frequency.

Those three, dated and repeated on a fixed cadence, are a measurement programme. Everything beyond them is elaboration.

Step 6: re-measure on a cadence, and after every change

Monthly suits most businesses; weekly suits fast-moving categories. Always re-measure after a change intended to affect visibility, because otherwise you have made a change and learned nothing.

Hold the prompts constant between rounds. Changing the prompts and the site at the same time makes the result uninterpretable.

What the numbers cannot tell you

They cannot tell you why an engine chose what it chose. A cited-source list shows what was used, not the reasoning. They cannot be projected forward — retrieval behaviour changes without notice. And they are not a ranking: a mention rate is an observation about a system that is free to answer differently tomorrow.

They also cannot be compared across tools that sample differently. A vendor reporting one run per prompt and a vendor reporting five are not producing comparable numbers, whatever both call the metric.

Doing this without a tool, and with one

You can do all six steps by hand. It costs nothing but time, and for a single business with a handful of prompts it is entirely reasonable — open each engine, run the prompts, keep a spreadsheet with the fields above.

It stops being reasonable at scale: four engines times ten prompts times five runs is 200 sessions per round, and the recording is where accuracy dies. That is what tooling automates. A VisiScan scan generates buyer-intent prompts for your category, runs them across up to four engines, and records every field above with quoted evidence, then audits your site for the readiness signals that explain the result. The methodology documents exactly how each observation is produced.

Related reading: how to get cited by ChatGPT, how to rank in Google AI Overviews, and GEO vs SEO vs AEO.

FAQ

How do you measure AI visibility?

Ask buyer-intent prompts across several AI engines, repeat each prompt several times, and record for every run whether your business was named, which competitors were named, and which sources were cited. Convert the runs into a mention rate, a competitor share, and a citation set, all dated.

How many times should each prompt be run?

Three to five runs per prompt per engine is usually enough to separate a stable result from noise. Fewer than three cannot distinguish a real signal from the model's own variance.

What is a good AI visibility score?

There is no absolute benchmark, because the number depends on your category, your prompts, and each engine's current behaviour. The useful comparison is against your own previous measurement and against the competitors named instead of you.

Should brand-name prompts be used?

Only as a secondary check. A brand prompt measures recognition; a buyer-intent prompt measures recommendation. Businesses routinely score well on the first and score zero on the second, and it is the second that produces customers.

How often should AI visibility be re-measured?

Monthly for most businesses, weekly in fast-moving categories, and always immediately after a change intended to affect visibility. Keep the prompts identical between rounds so the comparison means something.

Can AI visibility be measured for free?

Yes, manually — open each engine, run your prompts, and record the results in a spreadsheet. The free AI visibility scanner automates a first pass, and the free crawler and schema checkers cover the readiness signals that most often explain a poor result.

Keep going

Turn the ideas in this article into a measurable baseline for your own site.