Direct answer: an AI search visibility audit is a repeatable observation study. Select a fixed set of questions tied to customer decisions, test them on the AI search surfaces your audience may use, and record the answer, brand mentions, cited URLs, competitor appearances, and factual errors. Repeat the same sample under documented conditions. Report each observation separately. There is no defensible universal AI rank, and a single citation does not prove authority, traffic, or a conversion.
This guide owns the measurement workflow. It does not repeat the page-eligibility and evidence-structure work in the AI citation guide, and it does not turn every impression or mention into a business outcome. For that broader scorecard, use the zero-click measurement guide.
What the audit can and cannot prove
| Observation | What it supports | What it does not prove |
|---|---|---|
| Your brand appears in an answer | The system associated the brand with that question in this run | That users saw the answer at scale or visited the site |
| Your URL is cited | The page was presented as a source for this response | A fixed rank, endorsement, or conversion |
| A competitor appears repeatedly | The competitor has a recurring association or source pattern worth investigating | That copying its page will reproduce the result |
| An owned page is never cited | A gap may exist in eligibility, relevance, evidence, or sampling | That the page is blocked or low quality |
| A platform report shows citations or impressions | The site had measured exposure on that platform under its counting rules | Equivalent exposure across ChatGPT, Google, Bing, or Perplexity |
Google states that AI Overviews and AI Mode can use different models and techniques, and that the displayed responses and links vary. Microsoft says its AI Performance citation totals do not indicate placement, authority, or the role of a page in an individual answer. Treat volatility as part of the evidence model, not as noise to hide.
1. Define the business decisions first
Start with the decisions a customer makes before contacting or buying from the business. A useful audit asks whether the brand is visible when someone defines a problem, compares approaches, builds a shortlist, checks fit, or looks for proof. It does not begin with hundreds of synthetic keyword variants.
Write a one-sentence audit job, such as: “Find where our brand, evidence, or pages appear when Philippine small-business owners compare practical technical SEO support.” Then set the country, language, audience, product or service boundary, named competitors, platforms, and test period.
Keep branded and non-branded prompts separate. A brand appearing for its own name is a different signal from appearing in a category shortlist or problem-solving answer. Combining them can make the result look healthier than it is.
2. Build a small, fixed prompt set
Use customer language from sales calls, support questions, Search Console queries, reviews, proposals, and on-site search. A practical first pass is five prompt families with three to five prompts each. This is a manageable sample, not an industry standard. Keep the exact wording unchanged for the baseline.
| Prompt family | Question it tests | Example pattern |
|---|---|---|
| Problem definition | Does the category appear when the customer names a symptom? | How should a small business diagnose [specific problem]? |
| Approach or solution | Which methods and source types are recommended? | What is the safest way to [complete task] for [business type]? |
| Comparison | Which options and tradeoffs enter the answer? | [Approach A] vs [approach B] for [constraint] |
| Shortlist and fit | Which brands or providers are named for a defined need? | Who can help with [job] in [market] for [business size]? |
| Proof and risk | Which evidence, limitations, and failure modes are cited? | What should I verify before choosing [service or tool]? |
Exclude prompts that no real customer would ask, prompts that merely restate your page titles, and broad questions with no business decision behind them. If local intent matters, specify the real location in the prompt and test context. Do not add a city solely to force a local answer.
3. Run controlled observations
Test the interfaces that matter to the audience. A small business may choose Google AI Mode or an AI Overview, ChatGPT Search, Microsoft Copilot Search, or Perplexity. Availability and behavior can differ by account, device, country, language, and date, so record the visible interface rather than assuming one generic “AI engine.”
- Use a fresh conversation or equivalent clean starting state for each prompt.
- Keep prompt wording, language, and location stable for the baseline.
- Record whether a search-grounded answer appeared. A normal model answer without visible web sources is a different condition.
- Save the date, platform, mode, account state, country, device, exact answer, source panel, and cited URLs.
- Open every citation that affects the audit. Confirm the destination resolves and supports the nearby claim.
- Repeat the same sample on a second date before treating an isolated result as a pattern.
Do not “improve” a disappointing prompt midway through the baseline. Put revised or conversational prompts into a separate exploratory set. This preserves a comparable fixed panel while still allowing research into how follow-up questions change the answer.
4. Capture an evidence ledger
One row should represent one prompt on one platform in one run. Store screenshots only when they can be kept without private account details. The structured text record is the durable source because interfaces change and screenshots are difficult to compare at scale.
Run date and time:
Platform and visible mode:
Country, language, device, account state:
Prompt ID and exact prompt:
Search-grounded answer shown: YES / NO
Brand mentioned: YES / NO / INCORRECT
Owned URL cited: URL / NONE
Citation supports nearby claim: YES / PARTLY / NO
Competitors mentioned:
Competitor URLs cited:
Wrong or outdated claims:
Answer or source snapshot location:
Reviewer and review date:
Decision: VERIFY / IMPROVE / MONITOR / IGNORE
"run_date_time","platform_and_visible_mode","country_language_device_account_state","prompt_id","exact_prompt","search_grounded_answer_shown","brand_mentioned","owned_url_cited","citation_supports_nearby_claim","competitors_mentioned","competitor_urls_cited","wrong_or_outdated_claims","answer_or_source_snapshot_location","reviewer","review_date","decision"
Separate a company name in the prose from an owned URL in the source list. Also separate a correct citation from a merely visible citation. An outdated page or irrelevant source can create visibility that needs correction rather than celebration.
5. Calculate transparent rates
Use rates that a reviewer can rebuild from the ledger. Always show the numerator, denominator, platform, and date range. “Eight owned-source citations across 60 search-grounded responses” is inspectable. “AI visibility score: 73” is not.
| Metric | Calculation | Use |
|---|---|---|
| Answer availability | Search-grounded answers / attempted prompt runs | Shows whether the surface actually answered the sample |
| Brand mention rate | Correct brand mentions / search-grounded answers | Tracks association without pretending every mention is sourced |
| Owned-source citation rate | Answers citing an owned URL / search-grounded answers | Tracks visible source inclusion |
| Citation support rate | Owned citations that support the nearby claim / reviewed owned citations | Separates useful attribution from weak or misleading citation |
| Competitor appearance rate | Answers naming each competitor / search-grounded answers | Finds repeated associations by prompt family |
| Factual error rate | Answers with a material wrong or outdated brand claim / reviewed brand mentions | Prioritizes corrections and source maintenance |
Report results by prompt family and platform before showing an overall total. A provider can be absent from broad educational prompts but strong in local shortlist prompts. That distinction produces a useful decision. One blended percentage hides it.
6. Compare source and competitor patterns
Do not stop at counting names. Open the sources that repeatedly support competitor appearances and classify why they may be useful to the answer:
- First-party proof: service details, methodology, original research, pricing, policies, or a case with inspectable constraints.
- Independent validation: directories, reviews, associations, customer references, or trusted publications.
- Topical utility: a calculator, checklist, comparison, dataset, or answer that resolves a narrow sub-question.
- Fresh factual source: current product, availability, regulatory, location, or operational information.
- Incidental mention: a source appears once with no recurring relationship to the prompt family.
Record patterns, not imitation instructions. A directory may be relevant because the query asks for providers. It does not mean every business needs dozens of directory profiles. A competitor's product page may be cited for a current limit. That does not mean its page structure is a universal template.
7. Route each gap to an action
| Observed pattern | Likely next check | Possible decision |
|---|---|---|
| Brand is absent, but competitors appear with strong relevant sources | Compare page job, evidence, independent validation, and market fit | Improve an existing page or build a missing proof asset |
| Brand is mentioned, but no owned URL is cited | Inspect which third-party sources support the mention and whether the owned page answers the same claim | Clarify the owned source, strengthen entity consistency, or monitor |
| Owned URL is cited for the wrong claim | Check visible wording, dates, canonicals, duplicate versions, and source context | Correct ambiguity, consolidate versions, and request recrawl where appropriate |
| Old or incorrect brand facts appear | Locate the outdated owned and third-party sources | Update facts, correct source records, and document the change |
| No AI answer appears for most runs | Check whether the platform normally serves an AI answer for this query and context | Ignore the sample or monitor. Do not manufacture content for a nonexistent surface |
| One isolated citation appears once | Repeat the fixed prompt on another date | Monitor until a pattern exists |
Before creating a new URL, check whether the intended answer already belongs on an existing page. Use the brand SERP audit for owned-asset and reputation consistency, and the technical SEO audit workflow when crawlability or canonical evidence points to a sitewide problem.
8. Combine manual observations with platform data
Manual prompt testing shows the answer and its visible sources. Platform reports can show broader site-level exposure, but their coverage and counting rules differ.
- Google Search Console: Google announced a dedicated Generative AI performance report for Search and Discover in June 2026. The rollout began with a subset of sites. When present, it reports impressions, pages, countries, devices for Search, and dates. If the report is absent, do not infer zero AI visibility. Google's general guidance also says AI-feature traffic remains included in the Web performance report.
- Bing Webmaster Tools: AI Performance is in public preview. Microsoft says it reports total citations, average cited pages, sampled grounding queries, page-level citation activity, and trends across supported AI experiences. It explicitly says these metrics do not indicate ranking, authority, or placement in a specific answer.
- OpenAI and Perplexity: visible citations and source lists can be inspected in search-grounded answers. OpenAI's ChatGPT Help Center says search responses may include citations and that Sources can show cited sources plus other relevant links. Its API web-search documentation separately distinguishes inline citations from the larger set of consulted sources. Perplexity states that its answers include links to original sources. These source mechanics do not give a site owner a complete impression report.
- Business records: separately track qualified leads, assisted conversions, branded demand, and customer-reported discovery. Do not assign a lead to AI search merely because a citation existed during the same month.
9. Set a cadence that preserves comparability
Run the full fixed panel monthly when AI visibility is commercially important or quarterly when it is exploratory. Keep most prompts stable so changes are interpretable. Add a small rotating set for new products, customer questions, or market events, but never blend those new prompts into the baseline without labeling the change.
Trigger an extra focused run after a major source correction, site migration, rebrand, new service launch, or material platform-report change. Record the intervention date. A before-and-after observation is still not proof that the intervention caused the answer change, but it is more useful than an undated screenshot.
What not to do
- Do not report one universal AI rank. Platforms, modes, responses, and source sets differ.
- Do not count an ordinary answer as a citation test. Confirm that web search or a visible source-grounded mode actually ran.
- Do not use only branded prompts. They measure recognition, not category discovery.
- Do not treat every citation as correct. Open the source and check the claim it is presented to support.
- Do not create one page for every missed prompt. Map the gap to an existing page job first.
- Do not copy competitor pages because they appeared. Identify the source role and build evidence the business genuinely owns.
- Do not promise visibility from schema, an AI file, or a tool score. Google says no special AI schema or machine-readable file is required for its AI search features.
- Do not confuse exposure with revenue. Keep mentions, citations, clicks, leads, and sales as separate evidence levels.
Audit review checklist
- The audit has a written market, audience, country, language, service, platform, and date boundary.
- The fixed prompt panel uses real customer language and separates branded from non-branded intent.
- Every result records the exact prompt, visible mode, answer state, mentions, citations, competitors, and material errors.
- Cited URLs were opened and checked against the nearby claim.
- Rates show their numerators, denominators, platform, prompt family, and date range.
- Platform reports are interpreted under their own counting rules and availability limits.
- Each gap routes to verify, improve, monitor, or ignore before a new content request is created.
- A second reviewer can reconstruct the result from the evidence ledger.
The useful outcome is not a flattering score. It is a short, reviewable list of source corrections, evidence gaps, page improvements, and observations that should remain on watch.
Evidence basis
Platform behavior and measurement documentation were checked on August 31, 2026. The prompt panel, evidence ledger, rates, and action-routing model are FloxoLab audit frameworks.
- Google Search Central: AI features and your website
- Google Search Central: Generative AI performance reports in Search Console
- Microsoft Bing: AI Performance in Bing Webmaster Tools
- Microsoft Bing: Copilot Search and cited sources
- OpenAI Docs: web search output, citations, and sources
- OpenAI Help Center: Searching the web with ChatGPT
- Perplexity Help Center: answers and original-source citations
Need an evidence-led AI visibility baseline?
FloxoLab can define the prompt panel, preserve the source evidence, and turn recurring gaps into a focused SEO backlog without inventing a magic score.
Explore the SEO audit