CITERA :: ARTICLE
AI visibility: build your benchmark
Measure AI mentions, citations, and recommendations with a repeatable benchmark. Connect the results to organic traffic, qualified leads, and revenue.
Author
Liam Karlsson
An AI search visibility benchmark measures how often your brand appears in a defined set of AI answers, which sources those answers cite, and whether the resulting visits turn into customers. Its value is a repeatable baseline you can use to decide what to improve.
This guide sets out Citera’s recommended measurement method, with a worked example and a weekly review process. The example is illustrative; it is not a measured cross-industry study or a customer result.
Define the buying journey
Start with one audience, one market, and one commercial goal. A Swedish driving school needs a different question set from a global software company. Mixing both into one score makes it harder to choose the right work.
Use questions from sales calls, search queries, support requests, and customer interviews. Include the decisions that happen before someone knows your name. Keep branded questions in a separate group so searches for your own company do not inflate your discovery score.
Discovery: Which driving schools offer intensive courses in my city?
Comparison: How do local driving schools compare on lesson availability and price?
Qualification: Can I book evening lessons with an English-speaking instructor?
Decision: What is included in the course, and how do I book?
Our suggested starting point is 20 to 30 questions per priority audience, divided across these intents. This is an operational starting point, not a statistically validated sample size. Add questions when they reveal a real buying need, and document each change.
Choose consistent conditions
Choose the AI products your buyers actually use. Record the product, search mode, visible model version when available, language, location, date, account state, and exact question. Test in a fresh conversation with consistent settings. Keep ChatGPT Search, Google AI Overviews, AI Mode, and standalone assistants separate.
Repeat each question on multiple dates. Three runs can expose obvious variation in an initial pilot, but they do not establish statistical confidence. Save the complete answer and citation URLs for every run. Do not assume a model API reproduces the consumer product’s answers.
Record a missing AI Overview as ‘no AI answer shown’, not a failed request. Keep that condition separate from timeouts, refusals, and collection errors. Count all attempted searches when reporting how often an AI feature appears; use successful AI answers for the answer-level metrics below.
Publish the question count, run count, exclusions, and measurement window next to your score. Compare like-for-like groups. When you change a prompt set or testing method, start a new series or rerun the old conditions alongside it.
Measure five distinct signals
Use explicit counting rules. Count a brand once per answer, even if it appears repeatedly. Resolve spelling variants to the same brand, but manually check ambiguous names. Preserve raw counts so a percentage can be checked.
Mention rate = valid answers naming your brand / all valid answers in the test x 100.
Owned-domain citation rate = valid answers linking to your website / all valid answers x 100. Track third-party pages about you separately.
Recommendation rate = valid answers positively recommending your brand / all valid answers x 100. A neutral mention or warning does not count.
Tracked-brand share of voice = your brand-answer mentions / all brand-answer mentions across the fixed competitor set x 100. This is not market share.
Factual accuracy = verified correct claims about your brand / all checked claims about your brand x 100. Mark unverifiable claims separately and report how many claims were checked.
A citation can appear without a recommendation, and a recommendation can appear without a link. Preserve those differences. Record list position as context only: the first brand named is not automatically the engine’s preferred option.
Read a worked example
Illustrative example: you ask 25 questions three times on one AI product and receive 75 valid answers. Your brand is mentioned in 18 answers, your website is cited in nine, and your brand is positively recommended in 12. The tracked competitor set receives 90 brand-answer mentions altogether.
Mention rate: 18 / 75 = 24%.
Owned-domain citation rate: 9 / 75 = 12%.
Recommendation rate: 12 / 75 = 16%.
Tracked-brand share of voice: 18 / 90 = 20%.
If the next matched measurement produces 24 mentions from 75 answers, mention rate rises from 24% to 32%: an eight-percentage-point increase, or roughly 33% relative growth. Report both the counts and the percentage-point change. That movement alone does not prove your edits caused it.
Break the result down by intent. Strong discovery coverage with weak qualification coverage suggests a different task from having no mentions anywhere. Prioritize the questions closest to a buying decision, then inspect the actual answers before assigning work.
Connect visibility to growth
Keep three views together: AI answers, website behavior, and commercial outcomes. For each priority landing page, track search clicks, identifiable AI referral visits, qualified enquiries, and attributed revenue. Agree the conversion definition and attribution model before comparing periods.
Google includes AI Overviews and AI Mode activity in Search Console’s Web performance data; it is not a clean standalone AI referral report. Google’s measurement guidance explains the scope. Avoid labelling every change in Google traffic as an AI result.
OpenAI documents the utm_source=chatgpt.com parameter on referral URLs. Use it alongside your analytics source data to inspect visits and conversions. OpenAI publisher guidance. Treat unattributed visits as unknown, not automatically AI-generated.
Bing’s AI Performance report provides citation activity for supported Microsoft AI experiences. It describes citations, not revenue, and its totals are not interchangeable with your own prompt panel. Bing AI Performance documentation.
Review matched periods and note seasonality, campaigns, product changes, and model changes. Where practical, compare edited pages with similar unchanged pages. A useful benchmark connects an observed gap to a testable improvement; it does not promise causation from a chart.
Turn gaps into useful work
No brand mentions: inspect which companies and source pages answer the question. Check whether your site clearly explains the relevant product, location, and use case.
Mentions but inaccurate facts: correct pricing, availability, locations, or product details on your own pages and relevant third-party profiles.
Competitors cited, your site absent: identify the missing evidence. Add a specific comparison, an original example, a documented result, or a clear answer where it belongs.
Citations but few visits: give readers a reason to continue, such as a useful template, detailed comparison, calculator, or product demonstration.
Visits but few enquiries: inspect message match, pricing clarity, customer proof, forms, and the next action on the landing page.
Make meaningful improvements to existing pages before creating multiple pages that answer the same question. Fix factual changes when they occur; review other priority pages on a sensible schedule. Rewriting a date is not evidence that the content improved.
Run a weekly review
Keep one row per question and run in your measurement sheet. Include the segment, engine, test conditions, answer, named brands, cited URLs, recommendation status, factual issues, and evidence link. In a separate action log, record the affected page, proposed change, owner, publish date, and outcome.
Review collection failures and unusual swings before interpreting scores.
Read the changed answers and check the pages they cite.
Choose a small set of improvements with a clear customer benefit.
Publish with the approval settings appropriate to the work.
Recheck the same questions and compare qualified traffic and conversions over time.
Citera brings monitoring and execution into that workflow: content creation, publishing, website improvements, distribution, and updates, with the controls your team chooses. Explore the product or read the State of AEO briefing for the wider strategy.
Common benchmark questions
What is a good AI visibility score?
There is no universal pass mark across different question sets. Compare the same audience, competitor set, language, engine, and conditions over time. A score is useful when it changes what you do and can be related to customer outcomes.
Does being mentioned mean we won the prompt?
Use a stricter definition for commercial decisions: a positive recommendation for the relevant need. Keep simple mentions, recommendation status, and answer position as separate fields so the result cannot be improved just by changing the definition.
Can this prove AEO caused revenue growth?
The benchmark alone cannot. It measures a defined sample of answers. Connect it to analytics and CRM data, use comparable periods, and document alternative explanations. Where attribution is uncertain, say so.
Put the next improvement into motion. Start for free to get started with Citera, or book a demo to review the workflow with our team.