An AI visibility audit measures where and how AI answer engines — ChatGPT, Gemini, Perplexity, and Claude — mention, cite, and represent your brand when buyers ask questions you should be winning. Run it once and you get a ranked gap list, a map of your top-cited pages, and a 30/60/90 remediation plan. The first run takes under an hour; after that, a regular check keeps your data current.
Your immediate next step: pick 10–20 buyer prompts and run them across all four engines. That single action produces more signal than any tool dashboard can give you without real query data.
Here’s what a complete audit delivers:
- A ranked list of citation gaps by engine and query type
- Your top-cited pages mapped to specific buyer prompts
- Accuracy and sentiment flags on branded responses
- A share-of-voice benchmark against category competitors
- A 30/60/90 remediation plan with owners and timelines
As Forbes notes, marketers who audit brand AI visibility by testing citation frequency, accuracy, and source pages find gaps they can actually prioritize. This guide walks you through every step. For context on how AI is reshaping local discovery, see how AI search is changing local business discovery and the local SEO optimization guide for service businesses.
Key Takeaways
A complete AI visibility audit measures citations, accuracy, sentiment, and share of voice across ChatGPT, Gemini, Perplexity, and Claude, then converts those findings into a prioritized 30/60/90 remediation plan you can run in 60 minutes.
| Point | Details |
|---|---|
| Start with 10–20 prompts | Mix branded, unbranded, and local intent queries across all four major AI engines. |
| Score every response | Use a 5-dimension rubric (mention, citation, accuracy, helpfulness, sentiment) to convert raw logs into ranked priorities. |
| Fix identity files first | Adding llms.txt and identity.json produces citation improvements faster than any other technical fix. |
| Track share of voice monthly | Calculate SOV by engine and compare against category competitors to measure remediation progress. |
| Trystellor automates the cycle | Weekly LLM tracking, 30 articles/month, and a 4,000-site backlink network replace five separate tools from $199/month. |
Table of Contents
- What should your AI visibility audit actually measure?
- How to build a prompt library that covers every buyer intent
- How to run the tests and capture results without losing data
- How to score branded and unbranded responses for citations and accuracy
- Which pages are AI systems citing, and are they technically ready?
- How to turn audit findings into a 30/60/90 action plan
- How to measure share of voice without naming competitors publicly
- Which tools should you use for manual vs. automated tracking?
- What a Trystellor AI visibility audit looks like in practice
- Compact audit checklist and time-and-cost comparison
- What marketers should watch next in AI visibility
- Trystellor replaces five tools with one platform for ongoing AI visibility
- Sources
- FAQ
What should your AI visibility audit actually measure?
Scope is where most audits fail. Teams either test too broadly and produce noise, or test only branded queries and miss the unbranded traffic that drives most new customer discovery.
Decide what you’re testing
Start with four decisions before you run a single prompt:
- Decide whether to audit the company name, product lines, or both.
- Determine if your market focus is national or local.
- Identify which service categories matter most for revenue.
- Choose whether to audit in English only or in multiple languages.
Choose your platforms
Test in this priority order: ChatGPT, Gemini, Perplexity, Claude. These four are commonly used AI-driven answer engines in the United States. If your industry has a vertical LLM (a legal research tool, a medical reference engine), add it after the core four.
Define your metrics
Search Engine Journal’s 15-question audit framework makes clear that visibility is not just about volume. You need to track six distinct signals:
- Mentions: Does your brand name appear in the response?
- Explicit citations: Is a URL from your domain included?
- Answer accuracy: Are the facts about your brand correct?
- Sentiment/tone: Is the framing positive, neutral, or negative?
- Topical association: Is your brand connected to the right category?
- Share of voice (SOV): What percentage of relevant answers include your brand vs. category competitors?
For every query you run, log: engine name, model version, query text, timestamp, raw response, citations present (yes/no, URL), and a brief accuracy note. That minimum data set is what makes your audit repeatable and auditable over time.
How to build a prompt library that covers every buyer intent
Use at least ten prompts to cover branded queries, unbranded category queries, and transactional intent variants so you see a broad picture of where your brand appears and where it doesn’t.
Prompt categories to include
- Branded informational: “What does [Brand] do?” / “Is [Brand] a reputable [service type]?”
- Branded transactional: “Should I hire [Brand] for [service]?” / “How much does [Brand] charge?”
- Unbranded category: “Best [service type] in [city]” / “Top-rated [service type] near me”
- Unbranded comparison: “Compare [service type] options in [city]”
- Problem-based: “Who can help me with [specific problem] in [location]?”
- Review-seeking: “What are people saying about [service type] providers in [city]?”
Sample prompts to start with
- “What is the best [your service] company in [your city]?”
- “Recommend a trusted [your service] provider near [your city]”
- “Who are the top [your service] businesses in [your state]?”
- “Is [Brand Name] a good choice for [service]?”
- “What do customers say about [Brand Name]?”
- “Compare [service type] options in [your city]”
- “How do I choose a [service type] provider?”
- “What should I look for in a [service type] company?”
- “Which [service type] companies have the best reviews in [your city]?”
- “What is the average cost of [service] in [your city]?”
Tag every prompt in a spreadsheet
Your prompt spreadsheet needs these columns: prompt text, intent type (branded/unbranded/comparison/problem), priority (high/medium/low), target landing page, engine results summary, citation flag (yes/no), and notes. This structure lets you sort by citation rate or accuracy failures and spot patterns fast. For competitive prompts, use category labels like “category leader” and “local alternatives” rather than actual competitor names in any document you share externally.
How to run the tests and capture results without losing data
A 30–60 minute timebox works for the initial run. Test ChatGPT first, then Gemini, Perplexity, and Claude in sequence. Each engine has a different retrieval architecture, so the same prompt will often produce meaningfully different citations.
Pro Tip: Use a fresh incognito window or a new API session for every engine. Chat history and session personalization can skew results — a logged-in session that has seen your brand before may over-cite it. For the most reproducible data, use API access with no system prompt rather than the consumer UI.
What to log for every response
- Engine name and model version (e.g., GPT-4o, Gemini 1.5 Pro)
- Exact prompt text
- Full raw response (copy/paste or screenshot)
- All URLs cited in the response
- Timestamp and session ID (or “incognito session #1”)
- A one-line accuracy note
Name your screenshot files consistently: [engine]_[prompt-id]_[YYYYMMDD].png. That naming convention makes it easy to compare the same prompt across engines and track changes over time without hunting through folders.
One practical note on privacy: avoid entering personally identifiable information into open chat sessions. Use sanitized prompts that describe your business category and location without including customer names, internal pricing, or proprietary data. This matters especially if your team is running audits for clients in regulated industries like healthcare or financial services, where audit governance frameworks recommend strict data handling controls throughout the process.
How to score branded and unbranded responses for citations and accuracy
Raw logs are just text until you apply a scoring rubric. A consistent rubric converts qualitative observations into numbers you can sort, trend, and present to stakeholders.
Scoring rubric
Score each response on five dimensions, each worth up to 2 points:
| Dimension | 0 | 1 | 2 |
|---|---|---|---|
| Mention | Brand not mentioned | Brand mentioned indirectly | Brand named explicitly |
| Citation | No URL present | Partial/indirect URL | Direct URL from your domain |
| Accuracy | Multiple factual errors | Minor errors or outdated info | Fully accurate |
| Helpfulness | Response doesn’t address query | Partially addresses query | Fully addresses query |
| Sentiment | Negative framing | Neutral | Positive or favorable framing |
A perfect score is 10. Responses scoring 7 or above are performing well. Scores below 5 signal a page or topic that needs immediate attention.
Categorizing accuracy failures
When a response contains errors, flag the error type: wrong facts (incorrect pricing, wrong location, wrong service description), outdated information (a service you no longer offer), or missing context (your brand is absent from a category where you should appear). The EDPB’s AI auditing checklist recommends combining documentation checks with adversarial testing to surface bias and accuracy gaps — the same logic applies here. Test edge cases, not just the easy branded queries.
Once you’ve scored all responses, aggregate by page URL. Which pages earn the most citations? Which pages are cited but contain factual errors the AI is repeating? That page-level view is where the remediation plan starts.
Which pages are AI systems citing, and are they technically ready?
Map every citation URL back to your site architecture. After scoring multiple prompts across the four engines, you will usually find that a small number of pages earn the majority of citations; these are your most valuable AI-facing assets.
Build your top-cited pages table
Technical readiness checklist for AI crawlers
AI systems can only cite pages they can read. The AI Visibility Checker validates the presence and consistency of AI Discovery Files (llms.txt, ai.txt, identity.json, brand.txt), checks crawler access rules for GPTBot, ClaudeBot, and PerplexityBot, and returns a deterministic score from 0–100 with file-by-file recommendations. Run it on your domain before you do anything else on the technical side.
Beyond that tool, check these items manually:
- robots.txt: Confirm GPTBot, ClaudeBot, and PerplexityBot are not blocked
- Schema markup: Service pages need
Service,LocalBusiness, orFAQPageschema from Schema - Canonicalization: Duplicate pages split citation signals; one canonical URL per topic
- Page speed and mobile: Slow or broken pages get deprioritized by AI crawlers
- Internal linking depth: Pages more than three clicks from the homepage are harder for crawlers to reach
The fastest fixes are identity files and schema. Adding llms.txt and identity.json can improve citation consistency within days. Structured data updates on your top-cited pages take a developer a few hours and produce measurable results within weeks. For a broader technical checklist that pairs with these AI-specific fixes, the essential SEO audit checklist for local service businesses covers the overlapping ground well.
How to turn audit findings into a 30/60/90 action plan
Findings without priorities are just a list of problems. The goal is a plan your team can actually execute, with owners and deadlines attached to every task.
Prioritization matrix
Sort every finding by two axes: impact (how much will fixing this improve citation rate or accuracy?) and effort (how many hours or dollars does it take?). High-impact, low-effort fixes go first.
Severity levels:
- Critical: Brand is absent or misrepresented in high-volume branded queries
- High: Top-cited page has factual errors or missing schema
- Medium: Unbranded category queries return no citation for your brand
- Low: Sentiment is neutral rather than positive on informational queries
30/60/90-day plan template
- Days 1–30 (Quick wins): Fix robots.txt crawler access, add llms.txt and identity.json, correct factual errors on top-cited pages, add FAQ schema to service pages, update outdated content on pages scoring below 5.
- Days 31–60 (Content and authority): Publish new pages targeting unbranded queries where your brand has zero citations, build backlinks to top-cited pages, expand internal linking from high-authority pages to underperforming ones.
- Days 61–90 (Measure and iterate): Re-run the full prompt library across all four engines, compare new scores against baseline, identify which fixes produced citation improvements, and update the prioritization matrix for the next cycle.
TrueFoundry’s audit cadence guidance recommends continuous evidence capture for access and enforcement issues, monthly reviews for trends, and quarterly reviews for vendor and permission changes. For AI visibility specifically, a monthly re-run of your core prompt library is a practical minimum. Weekly automated tracking is better if your category is competitive.
Re-run the audit after any major content change, schema update, or backlink campaign. Citation changes from structured data fixes can appear within weeks. Identity file changes often take effect faster.

How to measure share of voice without naming competitors publicly
Share of voice (SOV) tells you what percentage of relevant AI answers include your brand. It’s the single most useful competitive metric in an LLM visibility audit.
How to calculate SOV
Run your 20 prompts across all four engines. For each response, record whether your brand is cited (1) or not (0). Do the same for each category competitor, using internal labels like “Category Leader A” and “Local Alternative B” in your tracking spreadsheet.
SOV formula: (Number of responses citing your brand) ÷ (Total responses tested) × 100
If your brand appears in 8 of 20 responses, your SOV is 40%. If the category leader appears in 15 of 20, their SOV is 75%. The gap — 35 percentage points — is your remediation target.
What SOV gaps tell you
A large SOV gap on unbranded queries usually means the category leader has more top-cited pages, stronger schema, or more authoritative backlinks pointing to the pages AI systems are reading. A gap on branded queries often signals an accuracy or identity problem: the AI has outdated or incomplete information about your brand.
For internal reporting, keep competitor names in a separate tab locked to your team. Public-facing reports should use category labels only. This protects sensitive competitive intelligence and keeps your public content focused on your own brand’s performance.
The CTAIO enterprise audit checklist recommends pre-deployment and production audit distinctions with clear ownership at each stage — the same principle applies to competitive benchmarking. Assign one person to own the SOV tracking spreadsheet and set a monthly update cadence.
Which tools should you use for manual vs. automated tracking?
The right choice depends on how often you need data and how many markets you’re tracking.
Tools for prompt testing and response capture:
- ChatGPT, Gemini, Perplexity, Claude: Direct UI testing for initial audits; API access for automated or bulk runs
- Ahrefs: Backlink analysis and content gap identification to understand why competitors earn more citations
- TripleWhale: Attribution and traffic analytics to correlate AI citation changes with actual traffic and revenue shifts
- AI Visibility Checker: Deterministic technical scan for AI Discovery Files and crawler access
- Google Sheets or Airtable: Prompt library, response log, and scoring rubric management
Cost and time estimates:
- Manual audit (one-time, 20 prompts, 4 engines): 1–2 hours, no tool cost beyond existing subscriptions
- Weekly manual cadence: 3–5 hours per week, significant analyst time
- Automated platform: $199/month and up, reduces weekly tracking to minutes of review
Pro Tip: Automate when you’re testing more than 20 prompts, tracking more than one market, or need weekly data. A manual check is fine for a quarterly baseline, but weekly competitive tracking by hand doesn’t scale past one or two markets without burning out your team.
Deliverable templates to build and keep on file: prompt spreadsheet, response log with scoring rubric, prioritization matrix, and a SOV tracking sheet. For teams that need web infrastructure to support AI crawler access, Wexla Websites builds and maintains sites for home service businesses with the technical structure AI crawlers expect.
HBR’s research on scaling AI work makes the governance point clearly: operational controls and documentation loops are what let teams run large-scale audits and maintain results over time. Without a documented process and assigned owners, audit findings tend to stall at the spreadsheet stage.
What a Trystellor AI visibility audit looks like in practice
Trystellor delivers a free AI Visibility Audit within 48 hours of a 15-minute onboarding setup. The audit covers citation status across 25 buyer prompts, a competitor SOV benchmark, a full technical site report, a backlink profile review, a Reddit opportunity map, and a custom 90-day action plan.
Dashboard metrics you’ll see
- Citation rate by engine (ChatGPT, Gemini, Perplexity, Claude)
- Top-cited pages ranked by citation count and accuracy score
- Accuracy failure rate by page and query type
- Identity consistency score across AI Discovery Files
- SOV vs. category competitors by prompt type
How the platform uses audit outputs
After onboarding, Trystellor runs weekly LLM tracking using the prompts your buyers actually use. Results feed directly into the content production queue: pages with low citation rates get new GEO and SEO-optimized articles targeting the exact queries they’re missing. The 4,000-site backlink network reinforces the pages that need authority. Technical audit findings surface in the dashboard with one-click fixes so your team doesn’t need a developer on retainer for routine schema and crawler access updates.
For teams in healthcare, legal, or financial services, a review workflow holds every published asset for approval before it goes live. The AI-powered SEO approach for local service visibility explains how the content and technical layers work together to build citation authority over a single quarter.
Compact audit checklist and time-and-cost comparison
Before you run your first audit, confirm you have these steps covered:
- [ ] Audit scope defined (brand, products, markets, languages)
- [ ] Platform list confirmed (ChatGPT, Gemini, Perplexity, Claude, plus verticals)
- [ ] Prompt library built (10–20 prompts, tagged by intent and target page)
- [ ] Logging template ready (engine, model, prompt, response, citations, timestamp)
- [ ] Scoring rubric set (5 dimensions, 0–2 points each)
- [ ] Technical checks scheduled (AI Visibility Checker, robots.txt, schema review)
- [ ] Prioritization matrix prepared (impact vs. effort, severity levels)
- [ ] Owners assigned for each fix category (marketing, SEO, dev)
Time and cost by audit type
Identity file fixes (llms.txt, identity.json) produce citation changes fastest, often within days of deployment. Schema updates on top-cited pages typically show results within weeks. Content and backlink campaigns take one quarter to move SOV meaningfully.
What marketers should watch next in AI visibility
The most underestimated shift in AI visibility right now is machine-readable identity. Most brands have spent years optimizing for human readers and Google’s crawler. They’ve written title tags, built backlinks, and published blog posts. Almost none of them have told AI systems who they are in a format those systems can actually parse.
The adoption of AI Discovery Files (llms.txt, ai.txt, identity.json, brand.txt) is accelerating. Major LLMs are updating their crawler access policies faster than most marketing teams are tracking. If your robots.txt is blocking GPTBot or ClaudeBot right now, you may not know it until you run a deterministic scan and see the score. That’s a fixable problem, but only if you’re checking regularly.
My practical advice: run a full audit quarterly, run a technical-only check monthly, and automate weekly prompt tracking if your category is competitive. Don’t let the audit become a one-time project. Citation landscapes shift as LLMs update their training data and retrieval logic. A brand that earned strong citations in January may find itself displaced by March if a competitor publishes more authoritative content or fixes their schema first. The teams that stay ahead are the ones that treat AI visibility as a recurring measurement discipline, not a one-time cleanup.

Trystellor replaces five tools with one platform for ongoing AI visibility
Running a manual audit once is useful. Running it every week, across four engines, with 25 prompts, while also publishing content and building backlinks, is a full-time job. Trystellor does all of it for $199 per month, replacing what most teams currently cobble together from five separate subscriptions.

The platform delivers weekly LLM tracking across ChatGPT, Gemini, Perplexity, and Claude, 30 GEO and SEO-optimized articles per month, a 4,000-site backlink network, weekly technical audits with one-click fixes, and daily Reddit opportunity maps. The free AI Visibility Audit arrives within 48 hours of a 15-minute setup and includes your citation status across 25 buyer prompts, a competitor SOV benchmark, and a custom 90-day plan. No credit card required for the three-day trial.
See the full platform at Trystellor’s product page and start your free audit today.
Sources
- AI Visibility Checker: Free AI Readiness Scan
- If AI Can’t Find You, Neither Can Your Customers: How To Audit Your Brand’s AI Visibility
- The AI Search Visibility Audit: 15 Questions Every CMO …
- The Complete AI Audit Checklist for 2026
- AI Audit: The 10-Step Enterprise Checklist | CTAIO
- AI Audit Checklist 2026: What to Review and When
- Getting AI to scale
- Checklist for AI Auditing
FAQ
What is an AI visibility audit?
An AI visibility audit measures how often and how accurately AI answer engines (ChatGPT, Gemini, Perplexity, Claude) mention and cite your brand in response to buyer queries. It produces a ranked gap list and a remediation plan.
How long does an AI visibility audit take?
The first manual audit takes 30–60 minutes using 10–20 prompts across four engines. Ongoing weekly tracking takes 3–5 hours manually or minutes with an automated platform like Trystellor.
Which metrics should I track in an LLM visibility audit?
Track six metrics: brand mentions, explicit URL citations, answer accuracy, sentiment/tone, topical association, and share of voice across engines. Share of voice is the most useful competitive signal.
How do I check AI citations for my brand?
Run your branded and unbranded prompts in ChatGPT, Gemini, Perplexity, and Claude using fresh incognito sessions. Log every response, flag URLs cited, and score accuracy using a 5-dimension rubric. The AI Visibility Checker handles the technical file validation separately.
How often should I re-run the audit?
Run a full prompt audit monthly if your category is competitive. Run a technical-only check (AI Discovery Files, robots.txt, schema) monthly. Automate weekly LLM tracking if you need to catch citation shifts before competitors do.

