Back to blog

AI Visibility

The GEO Audit: What It Should Actually Find

A GEO audit that only confirms an engine knows your name is a wasted afternoon. Here is what a useful one measures, the prompt set it needs, and how to turn the findings into work your team can ship this week.

Jake ChurcherJake Churcher··11 min read
The GEO Audit: What It Should Actually Find
On this page

US searches for the exact phrase "geo audit" held at 320 a month from March through May 2026, more than triple January's 90, before easing to 260 in June (DataForSEO Google Ads keyword data, US). The term is new enough that most of the pages ranking for it are free browser-based tools and small GEO agencies rather than an established playbook, which is a problem, because a GEO audit run badly produces a long list of gaps and no order to work through them.

Key takeaways

PointDetails
An audit is a snapshot, not a crawlIt measures what buyers get when they ask an engine, not whether a bot can technically reach your pages
A brand can appear and still loseShowing up in most answers is compatible with losing every high-intent comparison prompt
Four things get measured togetherRecommendation status, share of voice, sentiment and citation quality, never one alone
The prompt set decides the resultGeneric prompts return generic findings; the set has to mirror the real buying journey
Every finding needs an ownerAn unassigned gap sits in the report until the next audit finds it again
The check works best as a loopQuarterly for strategy, weekly for the handful of prompts closest to a purchase

What does a GEO audit actually measure?

A GEO audit is a structured read of how AI answer engines currently handle a brand: which buyer questions surface it, which omit it, how it gets described when it does appear, and which sources the engines lean on to decide. That last part is what separates it from a technical crawl. A site can be perfectly indexable and still lose every comparison prompt to a competitor with a stronger third-party trail.

The useful version is scoped to the questions that sit close to a purchase decision, not to the whole category. A general sweep across a hundred loosely related prompts returns a long list of gaps with no order to work through, which is exactly how audits end up filed and forgotten rather than acted on.

Where a conventional SEO audit stops

Site architecture, technical health, backlinks and page speed still matter, and a conventional SEO audit reports on all of them: rankings, impressions, links, crawl errors. None of that says what a buyer receives when an engine compresses a market into three sentences and names one option.

PerplexityExample answer

Buyer asks: “best crm for a growing sales team

For a growing sales team, Pipedrive and HubSpot are commonly recommended for their ease of use and scalability. Pipedrive suits teams that want a visual, sales-first pipeline with minimal setup, while HubSpot fits teams planning to add marketing and service tools later.

An SEO audit can confirm this page ranks. It cannot tell you whether the engine that answered this question named your brand, named a competitor instead, or described your product using language from a two-year-old review.

AI engines synthesise answers from product pages, review platforms, comparison content, community threads and structured data, and they weight those sources differently to how a ranking algorithm does. A page ranking third on Google can still lose the citation to a Reddit thread that answers the question more directly, and a page that never ranks at all can get quoted whole if it is the clearest source an engine finds.

Build the prompt set from buyer intent, not vanity questions

Generic prompts produce generic findings. Asking "what is [company]?" confirms an engine recognises the name. It says nothing about whether the brand can win new demand, because nobody searches that way once they are genuinely deciding between options.

Start from the questions that map to the sales motion instead:

Prompt typeExampleFunnel stage
Category"best [category] for a [team type]"Awareness
Pain point"how do I fix [specific problem]"Awareness
Job to be done"software for [specific workflow]"Consideration
Vertical or use case"[category] tool for [industry]"Consideration
Alternatives"alternatives to [competitor]"Evaluation
Direct comparison"[brand A] vs [brand B]"Evaluation
Suitability"is [product] suitable for [company size]"Decision

The right mix depends on the market. A security SaaS needs trust, deployment and compliance questions in the set; a product-led collaboration tool needs workflow and team-size comparisons instead. Koalr runs a fixed set of 25 real buyer prompts across all seven tracked engines, ChatGPT, Claude, Gemini, Perplexity, Microsoft Copilot, Google AI Overviews and Google AI Mode, and building a defensible prompt set covers the mechanics of putting one together from scratch.

Four things to check for every prompt

A single presence check is a weak signal on its own. A brand mentioned once, in passing, while a competitor gets called "the best choice for mid-market teams" is not visibility worth reporting to leadership. A useful audit measures four things together for every prompt:

  1. Recommendation status. Whether the engine names the brand as the answer, mentions it as one option among several, describes it inaccurately, or leaves it out. The four states an AI answer actually produces covers this in full, and it is the axis most audits collapse into a single number.
  2. Share of voice. How often the brand appears against how often competitors do, across the same prompt set and the same window.
  3. Sentiment. The language an engine uses when it does mention the brand, since a neutral aside and a direct recommendation read identically on a presence-only dashboard.
  4. Citation quality. Which sources the engine is actually citing to back its answer. Which domains earn the citation breaks down what a strong versus a weak source looks like.

Read the results like a revenue-risk report

A headline score tells leadership whether the brand is gaining or losing ground. The prompt-level detail is where marketing and product teams find the actual work. Four patterns are worth checking for specifically:

PatternWhat it looks likeWhy it matters
Competitor displacementThe same rival gets named whenever buyers ask about a capability the brand genuinely offersA capability gap in the answer, not necessarily in the product
Positioning driftThe engine describes the product using old, incomplete or wrong category languageThe source material an engine is pulling from is stale
Citation weaknessThe engine relies on thin third-party sources or a dated review instead of the brand's strongest evidenceThe best evidence exists but is not the thing getting retrieved
Agent-readiness failuresPricing, use cases or core claims are fragmented across pages, so an engine cannot assemble a clean answerThe content exists but is not shaped for retrieval

These are not equal problems. A factual error about a core product usually deserves an immediate fix. A missing mention on a low-volume, low-intent prompt usually does not, and treating every gap as equally urgent is how an audit turns into a backlog nobody clears.

Match the fix to the failure

The instinct after finding a gap is to publish something quickly, and the wrong version of that instinct is another page that restates what the brand already says elsewhere. The fix depends on which of the four patterns above produced the gap.

Match the fix to the failure

  • Absent from high-intent prompts: state the category, buyer and differentiator in direct language on the page an engine is most likely to retrieve
  • Misrepresented: update the source material engines encounter first, specific claims beat generic ones
  • Losing comparison prompts: build a fair, genuinely useful comparison, not a thin page that only names the brand
  • Weak citations: strengthen third-party proof, case studies, reviews and documentation an engine can verify
  • Fragmented facts: consolidate pricing, use cases and core claims onto pages an engine can quote whole

See where your brand shows up in AI search

Enter your website and get your AI Visibility Score across 7 engines. First score in under 5 minutes, no setup.

Treat a GEO audit as an ongoing channel

AI answers change as models, indexes, citations and competitor content change. A quarterly snapshot is useful for setting strategy. It is too slow for a category where a rival can start winning a comparison prompt inside a month.

Set a working cadence instead: review the prompt-level detail, pick the highest-impact gaps, assign an owner, ship the fix, then check the same prompts again. Perfect coverage across every prompt that could exist is not the goal. Becoming the default answer on the prompts that create pipeline is.

A practitioner's view on running these

The tension every GEO audit runs into is coverage against action. Run the full 25-prompt set across seven engines and the report is complete and nobody works through it in one sitting. Run five prompts and the check is fast but risks missing the one where a competitor just started winning. The rule I use with clients: audit broad once a quarter, then pull five to eight competitor-alternative and comparison prompts into a shorter weekly check, because those are the prompts where a shortlist forms. Everything else can wait for the next full pass.

The uncomfortable part to admit is that fixing a citation gap does not move the score on a schedule anyone controls. I have watched a rewritten comparison page sit unchanged in an engine's citations for six weeks before Perplexity picked it up, while a client reasonably wanted to know why the dashboard had not moved. The honest answer is that retrieval refreshes on the engine's timeline, not the team's, and an audit tells you what to fix, not when the fix will show up in the answer.

Where Koalr fits

Koalr's GEO Audit runs the same four checks this article describes, recommendation status, share of voice, sentiment and citation quality, against 25 real buyer prompts across all seven tracked engines, and rolls the result into an AI Visibility Score from 0 to 100 so the prompt-level detail and the headline number stay connected instead of drifting apart. Off-page sources that convert into a citation covers the fix for the citation-quality axis specifically, which is usually the slowest of the four to move.

Frequently asked questions

Frequently asked questions

A structured review of how AI answer engines currently handle a brand: which buyer prompts surface it, which omit it, how it gets described, and which sources the engines cite to decide. It measures the answer layer, not just whether a page is technically crawlable.

Sources

Written by

Jake Churcher

Jake Churcher

Co-founder, Koalr

Jake is co-founder of Koalr, where he works on GEO and AXO: how brands get found, cited, and recommended by AI. A decade in IT productising and marketing services, now applied to AI search. Passionate about building AI tools that help businesses win.

See how your brand shows up in AI search.

Enter your website and get your AI Visibility Score across the engines your buyers actually ask. First score in under 5 minutes, zero setup.

Koalr
© 2026 Koalr Ltd. All rights reserved.