Audit methodology
What Is a ChatGPT Visibility Audit?
See what a reliable ChatGPT visibility audit should test, store, calculate, and disclose before you use the results to guide marketing decisions.
Updated 2026-09-17 · 8 minute read
A visibility audit tests real buyer conversations
A ChatGPT visibility audit asks a standardized set of customer questions and records how the business appears in the answers. The questions should be tailored to the business while following a consistent intent framework, allowing different audits to remain comparable without forcing every industry into the same wording.
For a focused audit, 15 locked questions run three times each, creating 45 planned observations. Repetition matters because AI answers are variable. The exact questions, model, timestamps, raw answers, and parsing version should be stored so the result can be reviewed later.
Neutral testing protects the result
The provider prompt should not secretly include the audited business name, website, competitor names, or an instruction to find the business. The business is identified after the response. This makes the result a test of natural visibility rather than a guided lookup.
If the approved customer question naturally includes the business name, that wording can remain. Otherwise, the provider receives only the customer question and permitted location context.
Coverage and failures must be transparent
A benchmark is only meaningful when the report states how many observations were planned, attempted, completed, failed, and excluded. A failed API request is not the same as a completed answer that omitted the business. Treating both as zero would understate visibility and hide an operational problem.
A reliable report also distinguishes a complete result from partial or insufficient data. Coverage rules should be versioned and applied consistently.
The report should show evidence, not just a score
A single score is useful for orientation, but the question matrix explains the result. Each row should show whether the business appeared, whether it was recommended, how often it appeared across the three runs, whether its domain was cited, and which alternative led when the business did not.
The score and diagnostics should be reproducible from stored observations. Recalculation must never overwrite the original provider responses or locked question set.