You run an AI search, see your company named with a citation, and assume your visibility work is paying off. Or a competitor appears first, so you assume it has won. Either conclusion can be wrong when it rests on one generated answer.
A useful AI search audit has to answer three separate questions: Is the claim correct? Does the cited page support it? Does the result persist when you repeat the search? Once you separate those questions, you can stop treating citations as proof and start measuring what users are actually likely to encounter.
Separate answer accuracy, citation support, and repeatability
An answer can be correct while citing the wrong page. It can also quote a page accurately even though the page itself contains an outdated or incorrect fact. A perfectly supported answer may disappear on the next run. These are different failures, and each requires a different fix.
| Layer | Question to ask | What a failure means | What you should do |
|---|---|---|---|
| Claim accuracy | Is the statement factually correct? | The model generated, repeated, or combined incorrect information. | Find the authoritative fact and identify where the wrong version may be coming from. |
| Citation support | Does the linked page substantiate the exact statement beside it? | The citation is related to the topic but does not entail the claim. | Record the mismatch and improve the page that should support the claim. |
| Source quality | Is the cited information current, specific, and appropriate for the claim? | The answer may be grounded in weak, stale, or indirect evidence. | Strengthen first-party evidence and correct external profiles you control. |
| Repeatability | Does the claim, citation, or recommendation recur across runs? | The observed result may be sampling variation rather than durable visibility. | Measure occurrence rates across repeated prompts and engines. |
A citation is reliable only when the linked material materially supports the claim attached to it. Topical relevance is not enough. A page about a business does not automatically support every statement an AI answer makes about that business. Authority does not repair that mismatch either: a respected domain can still be the wrong citation for a particular sentence.
This is why accuracy belongs at the claim level. Work involving 158,000 AI claims validated through FactCheck used individual claims as the unit of analysis rather than assigning one broad true-or-false label to an entire response. Your audit should use the same basic unit. One answer may contain several supported claims, one unsupported inference, and one factual error.
Audit each AI answer at the claim level

Start with the exact answer the user saw. Do not rewrite it into a cleaner version before checking it. Small qualifiers such as location, availability, price conditions, service area, or timing often determine whether a citation really supports the statement.
- Capture the query context. Save the precise prompt, AI product or search surface, displayed model when available, location, date, and whether the session was signed in or personalized. A later result is not comparable if those conditions changed.
- Split the answer into atomic claims. Turn “Company A offers emergency plumbing throughout Toronto and is open all night” into separate claims about the service, service area, and hours. A citation may support one part without supporting the others.
- Mark opinions separately. Statements such as “best,” “most reliable,” or “ideal for families” are conclusions, not simple facts. Identify the factual premises that would be needed to justify the conclusion.
- Open every cited URL. Find the passage, field, table, or listing that is supposed to support the claim. Do not give credit merely because the page mentions the same entity or topic.
- Score correctness and support independently. Verify whether the claim is true, then decide whether the cited page proves it. A correct claim with an unrelated citation is still a citation failure.
- Save a short evidence note. Record what the page supports, what it omits, and any conflicting detail. This makes later reviews possible even if the page changes.
Use a small, explicit verdict set so different reviewers make comparable decisions:
- Supported: The cited material clearly substantiates the entire claim, including its qualifiers.
- Partially supported: The citation proves only part of a compound claim or leaves an important qualifier unresolved.
- Unsupported: The page is related but contains no evidence for the claim.
- Contradicted: The cited material states something incompatible with the answer.
- Unverifiable: The page is unavailable, the relevant content has changed, or the claim cannot be checked from accessible evidence.
Do not let a polished sentence hide a weak inference. If an AI answer calls a provider “the best option” because it has evening hours, the hours may be supported while the recommendation is not. Record the factual premise as supported and the superlative as unsubstantiated unless the answer supplies a defensible comparison.
The resulting audit should preserve four separate fields: the claim, its factual verdict, its citation-support verdict, and the reason for each verdict. A single “accurate” column collapses too much information to guide a correction.
Measure AI visibility as a distribution, not a ranking

Traditional rank tracking encourages you to ask where a business appeared. Generative search requires an earlier question: how often did it appear at all?
The instability can be substantial. Across 14,472 Gemini citations from 1,487 local queries in 50 large U.S. metro areas and ten service categories, repeated identical searches produced only about 40% overlap among cited sources. Gemini selected the same top business about 7% of the time, while a Google local-pack control returned the same top listing about 90% of the time.
Engine-to-engine agreement was even lower in that local-search sample. Gemini and ChatGPT cited the same domains in only about 8% of the compared searches and recommended the same top business 4.2% of the time. Gemini leaned heavily on business websites, while ChatGPT relied more on Reddit and business directories. Success in one engine therefore cannot stand in for visibility across AI search as a whole.
Those percentages are not universal benchmarks. They come from a defined set of U.S. local-service searches and should not be projected onto every industry, country, prompt type, or AI product. They do establish why a screenshot from one run is weak evidence of either success or failure.
A practical starter protocol, rather than a claim of statistical certainty, is to select ten commercially important prompts and run each one five times per engine. Keep the wording and observation conditions fixed. Treat alternative phrasings as separate prompts instead of changing the text between repetitions.
- Choose prompts by user decision. Include discovery, comparison, eligibility, trust, and branded-fact questions that can influence whether someone contacts or excludes you.
- Run a fixed batch. Capture every answer, including runs where your brand is absent and runs with no citation.
- Keep engines separate. Report Gemini, ChatGPT, and any other surface independently before creating an aggregate view.
- Repeat on a consistent cadence. Use the same batch before and after material content changes, and maintain unchanged prompts as controls.
- Compare rates, not anecdotes. Look for changes across the batch rather than celebrating or diagnosing one favorable result.
Calculate at least four rates:
- Mention rate: Runs that mention your entity divided by all runs for that prompt and engine.
- Citation rate: Runs that cite your domain divided by all runs.
- Recommendation rate: Runs that recommend your entity, with a separate field for first or primary recommendation.
- Supported-citation rate: Audited citation occurrences that fully support the attached claim divided by all audited citation occurrences.
Do not report “average rank” without a written rule for absent brands, unordered lists, and narrative recommendations. In many generated answers, numerical position implies a precision the interface does not provide. Mention and recommendation rates are usually easier to interpret.
This approach also prevents you from mistaking normal variation for the effect of an optimization change. If visibility rises from one run to the next while unchanged control prompts move just as much, you do not yet have convincing evidence that your edit caused the difference.
Build pages that can support the claims you want cited
Your own website is not merely a conversion destination. It can be the evidence layer behind an AI answer. In the defined Gemini local-search sample, nearly 60% of citations led directly to business websites, more than the combined share for directories, review platforms, and forums. Reddit was the second-largest category at 13.7%.
That does not mean publishing a page guarantees selection. It means you should give an AI system a clear, defensible first-party page to cite when it needs to verify a claim about you.
Create a claim-to-page map
List the claims that matter in a buying decision, then assign one canonical page to substantiate each one. Typical groups include services offered, locations served, eligibility or customer fit, operating hours, pricing conditions, product capabilities, policies, credentials, and named people responsible for the work.
For every claim, ask:
- Is the answer stated directly in visible page copy?
- Does the page identify the exact company, product, service, and location involved?
- Are conditions and exclusions placed beside the claim rather than hidden elsewhere?
- Does the page contain evidence appropriate to the statement?
- Is there a clear owner responsible for keeping the fact current?
- Does the page use a stable canonical URL that can remain valid when the content is updated?
A vague marketing page forces the answer engine to infer. A factual page reduces the number of inferences it has to make. Replace “solutions for every need” with explicit services, intended users, locations, and constraints. If availability depends on location or plan level, state that condition in the same passage.
Make JSON-LD agree with the visible evidence
Treat structured data as a machine-readable map of facts that a person can also verify on the page. For a local organization, use the most specific applicable Organization or LocalBusiness type and populate relevant properties such as name, URL, telephone, address, opening hours, and service area only when the page substantiates them.
Do not use JSON-LD to introduce claims the visible content cannot support. If the markup says a location is open all night but the location page lists limited hours, you have created ambiguity rather than authority. The same rule applies to ratings, prices, service areas, authors, dates, and product availability.
Check consistency across the page title, headings, body copy, structured data, internal links, and canonical URL. Schema cannot rescue a fact that is vague, contradictory, or attached to the wrong entity.
Audit external descriptions without manufacturing consensus
Your website may dominate citations in one engine while community discussions and directories carry more weight in another. Search for your brand, products, locations, and key claims across the pages that already appear in AI answers. Flag incorrect hours, old service descriptions, duplicate listings, former locations, and unsupported reputation claims.
Correct profiles and listings you legitimately control. Where a third-party page has a documented correction process, submit accurate evidence. Do not create fake reviews, staged forum discussions, or undisclosed endorsements to imitate independent agreement. Apart from the ethical problem, manufactured material gives answer engines more low-quality claims to misread and repeat.
When an inaccurate AI claim recurs, trace the wording across cited and uncited pages. If several pages repeat the same obsolete fact, updating only your homepage may not resolve the conflict. Record which representations you control, which have correction channels, and which must simply be monitored.
Key takeaways
- A correct answer can still have an unreliable citation, so score factual accuracy and citation support separately.
- Audit atomic claims, not entire responses. Compound sentences often mix supported facts with unsupported conclusions.
- One AI result is an observation, not a visibility trend. Repeat identical prompts and report occurrence rates by engine.
- Do not assume visibility transfers between Gemini, ChatGPT, or other AI search surfaces; their source preferences and recommendations can differ sharply.
- Publish canonical factual pages, align their visible content with JSON-LD, and correct external descriptions you legitimately control.
- Judge optimization work by changes across a fixed prompt set, not by a favorable screenshot.
On your next monitoring pass, keep the first batch deliberately small: ten decision-stage prompts, five identical runs per engine, and a claim-level review of every citation. That baseline will show whether your immediate problem is inaccurate information, weak evidence, unstable visibility, or a combination of all three. Fix the diagnosed layer, then rerun the same batch before expanding the program.
References
- Try Profound Blog – Where do inaccurate AI claims come from?
- Search Engine Land – Business websites dominate Gemini local AI search citations: Study






















