A·KAICiteKitGENERATIVE ENGINE RESEARCH
All posts
·AICiteKit Team

AI Search Audit Checklist: 30 Checks for Prompts, Sources, and Claims

A practical AI Search audit checklist for reviewing prompts, answer captures, citations, source quality, factual accuracy, and measurement boundaries before publishing a GEO report.

#ai-search#geo#ai-citations#measurement#content-strategy

The short answer

An AI Search audit is a documented review of a defined prompt sample, the answers it produced, the sources shown with those answers, and the claims a team wants to make about them. It is not a single score and it is not a shortcut to a universal AI ranking.

Use this sequence:

scope → prompt → answer → source → claim → action → retest

A good audit answers four questions:

  1. What exactly was tested?
  2. What did the answer actually say and cite?
  3. Which claims are supported, uncertain, or contradicted?
  4. What is the smallest defensible next action?
AI Search audit workflow connecting prompt design, answer capture, source review, claim verification, and a bounded action
Keep prompt context, answer evidence, source review, and the resulting decision in one auditable chain.

This checklist targets practical searches such as “how to audit AI Search visibility,” “what should a GEO audit include?”, and “how do I verify AI citations?” It is an evidence framework, not a claim that any edit guarantees a mention, citation, recommendation, click, or revenue outcome.

What an AI Search audit can and cannot establish

Google’s documentation says AI features in Search use the same foundational SEO requirements as Search generally and that appearing in AI features is not guaranteed (Google Search Central: AI features and your website, checked September 16, 2026). This supports an accessibility and content-quality baseline. It does not publish a universal citation formula.

OpenAI documents web search as a tool that can return citations and sources (OpenAI Developer Docs: Web search, checked September 16, 2026). That describes a documented OpenAI implementation; it does not make results from every AI surface interchangeable.

The independent research paper GEO: Generative Engine Optimization evaluates visibility in generative-engine responses (Aggarwal et al., arXiv, checked September 16, 2026). It is useful research context, not proof that a commercial metric predicts business results.

Audit observation It can support It cannot prove by itself
Brand appears in a captured answer Presence in that defined sample Universal market visibility
Brand is recommended Recommendation language in that answer Product quality or buyer satisfaction
URL is cited A source appeared with the answer A click, visit, or conversion
Crawler requests a page An identified request in server logs That the page was used in an answer
Answer changes after an edit A before/after pattern under recorded conditions That the edit caused the change

The 30-point AI Search audit checklist

A. Scope and decision: checks 1–5

  • 1. Name the decision. Are you checking factual accuracy, category discovery, source coverage, competitor representation, or traffic attribution?
  • 2. Define the unit of observation. Is one row one prompt run, one answer, one cited URL, or one claim-to-source relationship?
  • 3. Set the market. Record country, language, city or region when relevant, and any personalization context that is available.
  • 4. Separate surfaces. Keep ChatGPT Search, Google AI features, Perplexity, Gemini, Claude, and other products separate unless an aggregate is explicitly labelled and reproducible.
  • 5. Set the evidence rule. Define what counts as a mention, recommendation, citation, source position, absence, failed run, and non-comparable run before collecting results.

Do not call a prompt panel “the market.” It is a sample designed for a decision. A small, stable panel can be more useful than a larger panel with mixed intent, markets, and definitions.

B. Prompt design: checks 6–10

  • 6. Assign a stable prompt ID. Use an ID such as CAT-014 rather than relying on a changing spreadsheet row.
  • 7. Preserve exact wording. “Best enterprise GEO platform” and “GEO platform for a two-person agency” are different questions.
  • 8. Label intent. Use consistent groups such as problem, category, comparison, branded, evaluation, local, or transactional.
  • 9. Version meaningful changes. A new market, language, surface, or materially different wording should create a new version.
  • 10. Keep branded and unbranded prompts separate. A product-fact query tests identity and accuracy; it does not test category discovery.

A prompt is part of the measurement contract. If it changes silently, a later “trend” may simply be a different question.

C. Answer capture: checks 11–15

  • 11. Save the collection date and time zone. Answers and source sets can change.
  • 12. Record the surface and mode. Record model or version only when the product discloses it; otherwise write not disclosed.
  • 13. Preserve the answer. Save a permitted export, transcript, or capture reference before reducing it to a metric.
  • 14. Separate mention, recommendation, and citation. A brand can be mentioned without an owned URL being cited.
  • 15. Store exact source URLs. A domain-only label loses the page-level evidence needed for review.

If a platform does not expose a field, write not available. Do not infer model routing, location, or source order from a dashboard label.

D. Source and claim review: checks 16–21

  • 16. Map each source to a claim. Record what sentence or product fact the URL appears to support.
  • 17. Classify provenance. Use first-party, independent, community, commercial comparison, vendor-selected customer evidence, or unclear.
  • 18. Check relevance. Does the source answer the prompt’s intent, or is it merely adjacent?
  • 19. Check factual accuracy. Compare product facts, prices, integrations, and limitations with current primary documentation.
  • 20. Check freshness. Record publication, update, or access dates when available; stale pages can still be indexable.
  • 21. Check bias and incentives. A competitor, affiliate, agency, or vendor-selected case study may be useful but should not be presented as neutral evidence.

A source can be relevant without being independent. A vendor case study can document what the vendor says happened, but it is vendor-selected customer evidence, not independent proof of a general result.

E. Technical and measurement checks: checks 22–26

  • 22. Check the canonical destination. Follow redirects and record whether the cited URL still resolves to the intended page.
  • 23. Check access separately. Review robots directives, response status, rendered content, and server logs as adjacent evidence—not as an answer transcript.
  • 24. Check the denominator. Report matched runs, failed runs, unavailable captures, and excluded prompts.
  • 25. Check parser or product changes. Ask whether the monitoring tool changed its citation definition, source extraction, or historical calculations.
  • 26. Reconcile with analytics carefully. A citation, crawler request, referral, and conversion are different events and require separate measurement.

Google’s Search documentation is useful for foundational technical guidance, but an indexed page is not guaranteed to appear in an AI answer. Likewise, a crawler request does not prove that a page was cited.

F. Interpretation and action: checks 27–30

  • 27. Label confidence. High confidence requires complete context and answer/source capture; a summary score alone is low confidence for diagnosis.
  • 28. State what remains unknown. Missing model, location, raw answer, or source relationship should lower confidence.
  • 29. Choose one bounded action. Correct one stale fact, clarify one comparison criterion, fix one access issue, or schedule a matched retest.
  • 30. Predefine the retest. Keep the prompt, market, surface, and collection rules stable enough to interpret the next sample.

An audit is successful when a reviewer can inspect the evidence and understand the limitation. It does not need to end with a positive visibility result.

A practical audit record

Use one row per answer or claim-to-source relationship:

observation_id:
prompt_id:
prompt_version:
intent_group:
exact_prompt:
surface_and_mode:
model_or_version:
market_and_language:
collected_at:
answer_capture:
mention_status:
recommendation_status:
cited_url:
claim_supported:
source_class:
source_relationship_or_bias:
accuracy_status:
confidence:
next_action:
retest_date:

A useful accuracy status is supported, outdated, unclear, or contradicted. A useful confidence scale is:

  • High: complete answer and source capture, stable context, and claim checked against a current source.
  • Medium: answer and context captured, but a key field or source relationship is incomplete.
  • Low: summary metric only, incomplete answer, unknown context, or one unreplicated observation.

Confidence describes the quality of the record. It does not describe the probability of a future citation or business outcome.

How to audit a reported AI Search score

When a vendor or agency gives you a visibility, share-of-voice, citation, or position report, ask for the measurement contract before interpreting the headline:

  1. Which prompts, versions, markets, languages, and surfaces were included?
  2. What is the exact denominator, and how are failures handled?
  3. Does “visibility” mean a mention, recommendation, source link, or another event?
  4. Are exact answer captures and URLs available?
  5. Are source position and domain-level visibility separated?
  6. Are branded prompts separated from category and comparison prompts?
  7. What happens when a parser or platform integration changes?
  8. Can the report distinguish a vendor-selected source from an independent one?
  9. Can the result be exported for a human accuracy review?
  10. Which claims are observations, and which are editorial interpretation?

Tools can reduce collection work. AI Search Console, Peec AI, Otterly.AI, PromptWatch, and Profound describe different workflows and should not be assumed to calculate identical metrics. During a trial, compare raw answer access, exact source URLs, prompt versioning, surface context, failed-run handling, and export fields—not only the dashboard score.

What to do when the audit finds a problem

If the answer contains an incorrect product fact

Verify the claim against the current first-party page, pricing or documentation page, and any relevant independent evidence. Correct the canonical source if it is inaccurate or ambiguous. Then record the change and retest the same prompt. Do not claim that the correction caused a later answer change unless the design supports that conclusion.

Capture the criteria used in the answer and the supporting sources. The right action may be a clearer comparison page, better documentation, or an independent source gap—not automatically more keywords. A recommendation is evidence about one answer, not a verdict on product quality.

If the owned page is absent from citations

Check whether the prompt is matched, whether the answer cited any sources, whether the page is accessible and current, and whether another URL from the same domain was used. One absent citation is a lead for investigation, not proof of a visibility loss.

If the report claims traffic or revenue impact

Request the separate analytics and attribution evidence. A citation can coexist with a referral, but the citation alone does not prove that the user clicked, converted, or generated revenue.

FAQ

How often should an AI Search audit be run?

Use a cadence that matches the decision and the volatility of the facts being checked. A stable weekly or monthly panel is easier to compare than frequent runs with changing prompts. Pricing, availability, and other high-risk facts may justify additional reviews.

Is an AI Search audit the same as an SEO audit?

No. They overlap on technical accessibility and content quality, but an AI Search audit also preserves answer-level observations, source URLs, recommendation language, and the evidence boundary between citations and visits. Neither audit guarantees visibility.

How many prompts should an audit include?

There is no universal number. Start with the intent groups that matter to the decision, keep the denominator explicit, and add prompts only when they represent a defined question. More prompts do not repair mixed definitions or missing answer captures.

Can I use a single AI visibility score?

You can report an aggregate if its prompts, surfaces, weighting, missing-data rules, and definitions are documented. For diagnosis, keep surface-level observations separate because an aggregate can hide disagreement between systems.

Does adding schema guarantee AI citations?

No. Structured data can help communicate page information when it is accurate and eligible, but the sources checked here do not establish a universal rule that adding schema guarantees an AI citation or recommendation.

Sources and verification

The following sources were checked on September 16, 2026:

Last reviewed: September 16, 2026. Data confidence: High for the linked official and academic source descriptions; medium for the checklist, which is AICiteKit editorial guidance and should be adapted to the fields each platform exposes.