A mention is not a customer: how to measure AI visibility

A practical guide to repeatable AI search measurements: relevant questions, full answer evidence, explicit citations, incomplete results and separate inquiry tracking.

The short answer

Measure business mentions and explicit website citations separately, repeat a fixed set of relevant questions and keep the full answers. Track inquiries and bookings in a separate log. None of those outcomes should silently stand in for another.

Start with a real business question

Choose a service and an area the company genuinely serves. “Who repairs furnaces in [service area]?” asks about discovery. “What services does [business name] offer?” checks known-brand information. Keep discovery and branded diagnostic questions separate so a prompted name does not inflate discovery visibility.

Freeze a small, relevant question set

For a pilot, agree ten questions and three repetitions per selected engine. Record the wording, business identity, location, engine, model/settings and analysis version before running them. Four engines produce 120 planned answers per phase. This is a practical pilot design, not a statistical guarantee.

Keep an evidence record for every attempt

A URL listed among retrieved background sources is not automatically an answer citation. Check the provider’s citation fields and the visible answer evidence. If the business name is generic, review ambiguity rather than counting every word match.

  • Exact question, run number and timestamp.
  • Provider, model and relevant search/location settings.
  • Full answer and the explicit citation URLs returned with it.
  • Failure details and retry history.
  • Whether the business was mentioned and whether its verified website was cited.

Report each engine on its own

A worked example—not a customer result: suppose a complete engine batch contains 30 answers, six mention the business and two cite its website. Report 6/30 mentions and 2/30 website citations for that engine. Do not label the result “eight leads,” blend it with another engine or imply that every consumer will see the same response.

Withhold incomplete measurements

If a batch is incomplete, report the missing data and investigate. A failed request is not an answer where the business was absent. Retain every attempt; avoid selectively rerunning only unfavorable answers until the report looks better.

Make one change, then repeat

After reviewing the baseline, choose one verified issue and document the approved change. Repeat the same protocol at the agreed follow-up. A changed model, question set or analysis method can invalidate the comparison and require a new baseline.

Keep a separate inquiry log

A customer saying “Google” does not tell you whether they used Maps, ads, ordinary results or an AI answer. Do not turn that uncertainty into precise AI revenue attribution.

  • Date and requested service.
  • How the customer says they found the business; unknown is a valid answer.
  • Whether the inquiry was answered.
  • Whether an appointment or job was booked.
  • Other changes such as advertising, seasonality or service availability.

Read the result as evidence, not proof of causation

Answer engines and search results change without your intervention. Before-and-after observations can guide the next decision, but this small pilot has no automated significance test or causal study. Publish the limitations alongside the result.

Sources and related reading

This guide explains our approach. It is not a report of measured customer improvements.

Start with a useful question.

Which service do you want more customers for—and how do they find you today?

Discuss a pilot