Skip to content

Blog

AI visibility benchmarks: how to measure and publish credible results

Mike Holp · Published · Updated · Reviewed · 4 min read

An AI visibility benchmark is a repeatable measurement of how often a brand appears, is recommended, and is cited in answers to a defined prompt set. A credible benchmark publishes its engines, prompts, dates, classifications, sample size, and limitations. Without that context, a percentage is only a marketing claim.

Short answer: Define the audience and prompts first, run the same sample across the same engines, preserve the raw answers and citations, compare competitors within that sample, and report changes with dates. Never present a benchmark as universal search share.

What a benchmark should answer

A useful benchmark answers four questions:

  1. Where are we visible? Which engines and intent categories mention or recommend the brand?
  2. Where are we absent? Which high-value prompts return competitors instead?
  3. What evidence is being used? Which first-party and third-party URLs are cited?
  4. Did the pattern change? What happened after a documented content or brand change?

These questions keep the benchmark tied to decisions rather than vanity scores. The AI visibility metrics guide explains the measurements that belong in the underlying dataset.

A transparent benchmark method

1. Define the scope

Name the business, audience, locations, products, competitors, engines, language, and date range. If a location or engine is unavailable, state that instead of silently substituting another sample.

2. Build the prompt set

Use a stable core of branded, category, local, comparison, and problem prompts. The AI visibility prompt library provides a starting framework. Remove prompts that do not represent a real customer decision.

3. Establish classification rules

Define mention, recommendation, citation, position, sentiment or framing, and competitor presence before reviewing results. Apply the same rules to every brand. Keep uncertain results marked as uncertain rather than forcing a binary answer.

4. Preserve the evidence

Save the exact prompt, engine, date, answer, source URLs, page titles, and classification. A published summary can be concise, but the underlying evidence should be available for review. Citation-level tracking is covered in this URL tracking guide.

5. Compare within the sample

Calculate rates only from the prompts and engines actually measured. Competitor share is meaningful as a within-sample comparison; it is not the same as total market share or organic search share.

6. Repeat after a documented change

Record what changed, when it changed, and which prompts it was intended to affect. Rerun the same core sample before expanding the methodology. If answers vary substantially, report a range or the number of repeated observations rather than a single overconfident result.

Benchmark report template

SectionInclude
ScopeBrand, audience, location, engines, language, dates
SamplePrompt count, intent mix, competitor set
VisibilityMention and recommendation observations
EvidenceCited URLs, source types, and supported claims
ComparisonCompetitor presence within the same sample
Change logContent or technical changes with dates
LimitationsMissing data, volatility, attribution, and sample boundaries
ActionsTwo or three next steps linked to evidence

How to publish a case study without overclaiming

If you have permission and data, publish the baseline, intervention, follow-up date, prompt set, and measured change. Explain what else changed during the period. If you do not have permission or sufficient data, publish the method as a benchmark framework instead of inventing an outcome.

Competitor platforms publish customer case studies with claimed changes, such as Scrunch's Stratabeat case study. Those are useful examples of format, but their results should not be generalized to every business or treated as independent evidence for VisiScan.

VisiScan's self-audit article is a transparent first-party example. A future independent case study should add permissioned customer data and a repeatable before-and-after design.

Benchmark mistakes to avoid

  • Calling a small prompt sample “the market.”
  • Mixing branded and unbranded visibility into one unexplained score.
  • Changing prompts between baseline and follow-up.
  • Reporting competitor percentages from different samples.
  • Claiming that a content edit caused a result without a time-stamped change log.
  • Publishing customer names, answers, or analytics without permission.
  • Presenting competitor case-study claims as independently verified facts.

FAQ

What is a good sample size for an AI visibility benchmark?

There is no universal number. Start with a focused set large enough to cover the important intents and small enough to review consistently. Report the exact prompt count and composition so readers can judge the result.

Is AI visibility benchmark data comparable to SEO rankings?

Not directly. SEO rankings and AI answers use different surfaces and measurement conventions. An AI benchmark describes observed answer-engine responses for a defined sample.

How often should an AI visibility benchmark be repeated?

Repeat it after meaningful changes and on a cadence that matches the business decision. Keep the core prompt set and methodology stable so the comparison remains interpretable.

Can I publish a benchmark without customer data?

Yes, publish a methodology, first-party audit, or anonymized dataset only when it is accurate and permitted. Do not label a framework or illustrative example as a customer result.

Sources

Keep going

Turn the ideas in this article into a measurable baseline for your own site.