Your page can rank well in Google and still be absent from the answer your buyer sees. Ahrefs found that only 38% of pages appearing in Google AI Overviews also ranked in the traditional top 10, down from 76% eight months earlier. Organic rank is still useful, but it can no longer stand in for AI visibility.
You need a test that shows where visibility breaks: whether an AI system retrieves your brand, mentions it, cites it, explains it correctly, places it on a shortlist, or recommends it. The framework below turns those separate outcomes into a prompt panel, a repeatable scorecard, and an experiment you can act on.
Start with the decision, not a visibility score
AI visibility is not a single event. Your brand can be cited without being recommended, mentioned without receiving a citation, or described accurately but placed behind competitors. Treating all three situations as visible conceals the problem you need to fix.
Separate each answer into five measurement states:
- Retrieval: the AI answer appears and has an opportunity to include your brand.
- Inclusion: your brand, product, or page is mentioned.
- Attribution: an owned URL or a third-party page about your brand is cited.
- Positioning: the answer gives your brand a particular order, category, use case, or authority level.
- Recommendation: the answer actively includes your brand in the decision set for the intended user.
This separation reflects how mention order, explanation depth, authority framing, and comparative positioning can each change the value of an appearance. Decide which state matters before collecting answers.
| Your objective | Prompt family to test | Primary measurement | Guardrail |
|---|---|---|---|
| Correct the brand narrative | Branded identity and validation prompts | Factual accuracy and explanation depth | Owned citation rate |
| Expand category discovery | Unbranded category and problem prompts | Brand mention rate | Competitive share of mentions |
| Enter the buyer’s shortlist | Alternative, comparison, and decision prompts | Recommendation rate and mention order | Accuracy of the stated use case |
| Become a cited evidence source | Informational and how-to prompts | Owned-domain citation rate | Relevance of the cited page |
Denominators matter, especially on search surfaces that do not generate an AI answer for every query. A missing AI Overview is not the same result as an AI Overview that appears but omits your brand. Track both:
- AI answer trigger rate = attempts that produced an AI answer divided by all attempts.
- Among-answer mention rate = rendered AI answers mentioning the brand divided by all rendered AI answers.
- End-to-end mention rate = attempts mentioning the brand divided by all attempts, including attempts without an AI answer.
Do not compress these outcomes into one proprietary visibility score. A composite can rise because branded prompts improved while the unbranded prompts that create new demand deteriorated. Show the component rates and their numerators so a change remains interpretable.
Build a prompt panel that can be rerun

A useful prompt panel is a measurement instrument, not a loose keyword list. Every prompt needs a defined intent, an eligible engine or surface, and a reason for being in the panel.
- Branded identity prompts test whether the system knows what the brand is, who it serves, and how it differs.
- Category prompts remove the brand name and test discovery for the problem or product class.
- Comparison prompts test alternatives, versus questions, and the attributes used to separate competitors.
- Decision prompts add a buyer constraint, such as audience, use case, risk, or required capability, and test whether the brand is recommended.
- Validation prompts test reputation, limitations, suitability, or factual claims that a buyer may check before acting.
Keep a stable core panel for trend reporting and a separate exploratory panel for new questions. If you rewrite, remove, or add core prompts, create a new panel version. Do not splice the results into the previous trend line as though the test stayed constant.
Run each target engine as its own surface. A first-month fictional-brand test covering 825 prompts and 15,835 answers found materially different behavior across ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, and Gemini. Google AI Mode was comparatively stable for branded questions, Perplexity surfaced new material quickly, ChatGPT recognition strengthened during the month, and Gemini produced substantial citation gaps. Because the brand was artificial and the observation window was short, those results are evidence that engines differ, not a permanent ranking of the engines.
Repeat the exact prompt rather than trusting one screenshot. SE Ranking observed that Google AI Mode overlapped with itself only 9.2% when the same query was run three times. Three runs will not eliminate uncertainty, but they provide a practical first check on whether an appearance is repeatable or incidental.
For every run, store:
- A permanent prompt ID, prompt family, and panel version.
- The exact prompt text without silent edits.
- The engine and specific surface, such as Google AI Mode or Google AI Overviews.
- The date, run number, locale, and any account or session conditions you can keep consistent.
- The complete answer, ordered brand mentions, cited URLs, and first cited URL.
- Whether your brand was recommended, how it was framed, and whether the description was accurate.
Use a fresh conversation for each conversational-engine run so earlier messages do not become an uncontrolled input. Run repetitions in the same measurement window, then rerun the complete batch on a fixed cadence. Weekly measurement can suit an active intervention; a monthly cadence may be enough for an established baseline. Consistency matters more than choosing an arbitrary universal interval.
Evaluate tracking tools against this test design. Familiar SEO integration can still leave you with narrow LLM coverage and no optimization workflow. Before committing to a platform, confirm that it covers your target surfaces, retains raw answers and cited URLs, distinguishes mentions from citations, preserves prompt versions, records repeated runs, and exports answer-level rows. A polished summary dashboard cannot compensate for missing evidence.
Score each answer without losing its context
Create one row per answer, not one row per prompt. Aggregating three runs before storing them destroys the variation you are trying to measure.
- Inclusion: record brand absent or present. Calculate mention rate separately for branded, category, comparison, decision, and validation prompts.
- Attribution: distinguish an owned-domain citation from a citation to an independent page about the brand. Then record whether the owned page was the first or main cited source. A third-party citation can improve brand exposure without giving your site attribution.
- Order and recommendation: record the brand’s position among listed options and whether the language explicitly recommends it. Do not treat a neutral appearance in a list as a recommendation.
- Explanation depth: apply a small internal rubric consistently. Score 0 for absent, 1 for a name-only list appearance, 2 for a short explanation containing one defined claim, and 3 for a substantive explanation covering the audience, use case, or reason to choose. This is an operational rubric, not an industry benchmark.
- Framing and accuracy: label the tone as positive, neutral, cautionary, or negative. Record authority labels such as leader, challenger, or niche option only when the answer actually uses that framing. Mark factual descriptions as correct, incomplete, or incorrect in a separate field.
- Stability: with three runs, report whether the brand appeared in none, one, two, or all three. Keep that distribution visible beside the average rate.
Mention order deserves its own field because people often accept the shortlist they receive. A Growth Memo and Citation Labs test found that 74% of users selected the AI system’s first suggestion, while 26% changed the order when they recognized a brand they trusted. First position can provide an advantage, but it does not erase brand recognition, explanation quality, or trust.
Accuracy is the non-negotiable guardrail. A confidently worded but false recommendation is not a visibility win. Keep inaccurate claims in the visibility totals so you do not hide the problem, but flag them separately and prioritize correction over reach.
Report each metric with its numerator and denominator. A percentage without the number of eligible answers conceals small samples, missing AI-answer triggers, and changes to the prompt mix. Break results down by engine, prompt family, branded versus unbranded intent, and run consistency before looking at an overall total.
Turn signal patterns into controlled content changes

Diagnose the gap before editing
The scorecard should point to a failure mode. It should not merely tell you that visibility is low.
| Observed pattern | Likely reading | Next test |
|---|---|---|
| Strong branded mentions, weak category mentions | The entity is recognized, but its association with the wider problem or category is weak. | Test a page that connects the brand clearly to the category, audience, and use cases. |
| Frequent mentions, few owned citations | The brand is known, but the main site is not being selected as evidence. | Consolidate definitive facts on an owned page and inspect which independent URLs are being cited instead. |
| Citations without recommendations | Your material is useful as evidence, but the brand’s decision position is unclear. | Test explicit audience fit, differentiators, selection criteria, and honest limitations. |
| Name-only appearances | The system has too little usable information for a deeper explanation. | Test one comprehensive page that answers what the brand is, who uses it, and how to choose it. |
| Top placement in only one run | The apparent lead may be output volatility rather than a stable gain. | Repeat the batch and report the run distribution instead of publishing the best screenshot. |
| Visibility on one engine only | The gain is surface-specific. | Inspect that engine’s citations and distribution path; do not describe the result as universal AI visibility. |
| Positive but inaccurate descriptions | Repeated claims are shaping the narrative without adequate verification. | Correct the canonical brand information and monitor the exact false claim across owned and independent pages. |
Identity pages can matter earlier than broad authority. In the fictional-brand experiment, an About page and a consolidated brand guide became frequent citations, while detailed guides, reviews, and comparison pages performed better than generic formats. For a legitimate brand, that makes an accurate entity page and decision-oriented content sensible hypotheses to test. It does not guarantee the same outcome in every category or engine.
Do not assume a topical cluster is itself an AI visibility signal. During the first month of the same artificial setup, a hub with 10 supporting pages earned no citations, while 30 shorter, repetitive pages collectively generated more than 1,800 citations. That result does not establish repetition as a durable content strategy. It shows that site architecture alone is not a treatment, volume can create exposure, and visibility is not proof that a claim has been rigorously verified.
For your site, give each supporting page a distinct job tied to a real prompt or decision. Measure which URL is cited. Remove or correct pages that merely repeat claims, especially when repetition could amplify an error.
Test one explanation at a time
Most AI visibility work is a structured before-and-after test, not a true randomized A/B test. Retrieval systems change, answers vary, and you do not control when every engine discovers a revision. You can still make the evidence more useful:
- Write a falsifiable hypothesis. For example, clarifying audience and category on the canonical brand page should increase explanation depth on branded identity prompts.
- Capture a triplicate baseline batch. If the three runs conflict sharply, repeat the baseline before changing the site.
- Make the smallest coherent intervention. Update the entity page, publish a comparison resource, or improve a specific claim set, but do not combine a redesign, a large publishing sprint, and a distribution campaign if you want to know what helped.
- Record the changed URLs, publication date, affected claims, internal links, and prompt families expected to move.
- Use discovery as the gate instead of assuming every engine follows the same calendar. Begin interpreting the post-change period only after the new or revised material appears in citations or is otherwise demonstrably available to the surface being tested.
- Rerun the same panel, engine mix, session setup, and scoring rules. Keep newly discovered prompts in the exploratory panel until the current test ends.
- Compare the target prompts with unaffected prompt families and competitor patterns. If every brand moves in the same direction, engine drift is a stronger explanation than your page change.
- Repeat the result in another scheduled window. Call a one-engine or one-run gain directional, not conclusive.
Describe a before-and-after movement as associated with the intervention unless you have stronger controls. That language is not timidity; it is an accurate reflection of a system whose retrieval, citations, and generated wording can all change outside your test.
Keep AI response metrics beside traditional SEO and business outcomes. Citations do not guarantee visits, and visits do not prove that the answer influenced a decision. Some ChatGPT journeys continue on Google as users verify what they were told, so direct AI referrals may miss part of the path. Compare AI visibility with organic landing-page activity, branded demand, qualified visits, and conversions, but do not assign causation merely because two lines moved together.
Key takeaways
- Choose the decision you need to make before choosing a visibility metric.
- Separate AI-answer triggers, mentions, owned citations, independent citations, mention order, recommendations, explanation depth, framing, accuracy, and stability.
- Keep branded, category, comparison, decision, and validation prompts in separate cohorts.
- Measure each engine and surface independently, and run the exact prompt three times as a practical volatility check.
- Store one row per answer with the raw response and cited URLs. Do not rely on a composite score or a selected screenshot.
- Diagnose the missing stage, change one coherent content element, wait for discovery, and rerun the versioned panel.
- Track AI visibility beside rankings, traffic, and conversions without treating any one of them as a substitute for the others.
Your first useful measurement system can be a spreadsheet: a stable core prompt panel, three runs per prompt, one row per answer, and one intervention tied to one failure mode. Automate it after the process can explain why a number moved. That is the point at which AI visibility becomes an operating metric instead of a collection of interesting screenshots.
References
- Try Profound – In-Depth Review of SE Visible: A Solid Tool with Limitations
- Search Engine Land – Can a Fictional Brand Outsmart AI? Our Surprising Experiment
- Search Engine Land – Mastering AI Search Visibility: Key Signals You Need to Know

Leave a Reply