Back to blog

AI Visibility

AI Visibility Score: What It Measures and How to Move It

One number tells you whether you have a problem. The prompt-level detail underneath tells you which page to write. Here is what goes into the score, why two tools disagree about yours, and which levers actually shift it.

Joe CreightonJoe Creighton··13 min read
AI Visibility Score: What It Measures and How to Move It
On this page

An AI visibility score is a single number, normally 0 to 100, that combines how often AI answer engines name your brand, how prominently, whether they cite your pages as the source, and how that compares to competitors on the same questions. A low score means buyers asking about your category are being shown someone else. The number itself is a smoke alarm. Everything useful lives in the prompt-level detail underneath it.

PerplexityExample answer

Buyer asks: “best AI visibility tools for a small marketing team

For smaller teams, the tools most often recommended are Koalr and two of its competitors, which cover multiple answer engines at accessible price points.

Sources: a rival's roundup post, a Reddit thread

Illustrative example. The brand is named, but every cited source belongs to someone else.

Your brand is named, so mention rate goes up. But not one of the cited sources is yours, so citation rate does not move, and the buyer's next click lands on a competitor's page. A score that counted only mentions would record this as a win. That gap between the two is the whole reason the composite exists.

What does an AI visibility score actually measure?

It measures four things that behave independently, then weights them into one number. Understanding which component is dragging is what tells you what to fix, because the fixes are unrelated to each other.

ComponentWhat it asksWhat a low reading means
Mention rateOn what share of tracked prompts does an engine name you at all?Buyers in your category are never shown you. A content problem.
ProminenceWhere in the answer, and how high in any list?You are an also-ran. A positioning and third-party-proof problem.
Citation rateDoes the engine link your page as its source?Someone else's page is describing you. An off-page problem.
Share of voiceHow does your mention rate compare to named rivals on identical prompts?Absolute numbers look fine while you lose ground. A competitive problem.

The four come apart in practice. A brand can hold a 40% mention rate with a 2% citation rate, which means engines know the name but learn about it entirely from other people's pages. A brand can hold a high citation rate with poor prominence, which means its pages are trusted as reference material while a competitor gets the recommendation.

Treating the composite as one dial hides that. This is why a score is a starting point for a diagnosis, not the diagnosis.

Which engines should the score cover?

Seven surfaces matter, and they disagree with each other enough that a score from any one of them tells you very little about the rest. Koalr tracks all seven: ChatGPT, Claude, Gemini, Perplexity, Microsoft Copilot, Google AI Overviews and Google AI Mode.

The disagreement is measurable. Across 295,746 citations we looked at which page type each engine prefers to cite:

Page typeChatGPTCopilotPerplexityGeminiAI ModeAI Overviews
Homepage34.1%30.0%10.9%5.6%8.2%9.3%
Product or feature page11.0%4.0%21.2%10.5%23.7%24.4%
Pricing page14.8%12.3%24.1%4.0%21.2%22.6%
Blog post5.2%12.1%17.3%12.8%24.0%26.1%
Question-style page7.8%3.9%19.8%7.2%27.2%29.8%

Read down the columns and the strategy separates cleanly. ChatGPT takes 34.1% of its citations from homepages and 5.2% from blog posts, so a blog-only programme is close to invisible there. Google's two surfaces take 51% of their citations from question-style pages and 50% from blog paths, so they reward exactly the content ChatGPT ignores.

A tool that only checks ChatGPT and reports a low score is telling you your homepage is weak. It is not telling you anything about your blog.

How is the score actually calculated?

Most vendors, us included, build the composite the same way: a base appearance rate, adjusted by where in the answer the mention lands, then modified by whether a citation was attached and what sentiment the surrounding language carries.

  1. 1

    Run a fixed prompt set

    The same real buyer questions every cycle, across every engine. Koalr runs 25 per scan. A changing prompt set makes the trend line meaningless, because you cannot tell a real move from a different question.

  2. 2

    Record presence, position and source

    Per prompt, per engine: were you named, how early, and which URL was cited. This per-answer record is the data. The score is just its summary.

  3. 3

    Weight and combine

    Appearance rate multiplied by a position weight, plus a citation bonus and a sentiment adjustment, normalised to 0 to 100 so it is comparable across categories.

Three methodology choices decide whether the resulting number is worth acting on.

Where the prompts came from. Prompts derived from real search and buyer behaviour produce scores that track commercial reality. Prompts invented by the vendor produce a score for questions nobody asks. This is the single largest source of variance between tools.

How many there are. Small prompt sets move with model sampling noise. Any single answer can differ between two runs an hour apart with no change to your site at all, so a handful of prompts gives you a number that jitters. Enough prompts, held constant, averages that out.

When it was measured. Engines update. A score from three months ago describes retrieval behaviour that may no longer exist. A score with no date attached is not a measurement.

Why do two tools give the same brand different scores?

Because they are not measuring the same thing, and none of the differences are bugs. Expect a spread of 15 to 60 percentage points on the same brand in the same week.

Cause of divergenceEffect on the number
Different prompt setsThe largest single cause. Different questions, different answers.
Different engine coverageA tool tracking three engines cannot see the four where you might be strong.
Different position weightingBeing named fifth counts almost as much as first in some rubrics, and far less in others.
Different citation handlingSome count a mention with no source as a full hit. Others discount it heavily.
Different sampling cadenceA daily average and a weekly spot check produce different numbers from identical behaviour.
Geography and languageUK-phrased prompts return different sources than US-phrased ones.

The practical consequence: never compare your score across tools, and never celebrate a jump that coincided with a change to the prompt set. A score is only meaningful against its own history, measured the same way. Comparing your Koalr score to a competitor's Peec score is comparing two different measurements that happen to share a scale.

See where your brand shows up in AI search

Enter your website and get your AI Visibility Score across 7 engines. First score in under 5 minutes, no setup.

How do you establish a baseline worth trending?

The order matters. Most teams start by running prompts, then discover the prompts were wrong and lose their history when they fix them.

StepWhat to doWhy this order
1Write down the two to four buyer types who actually decidePrompt phrasing follows from who is asking
2Map the questions each asks by purchase stageAwareness, comparison and evaluation questions behave differently
3Fix the prompt set and version itChanging prompts later resets your trend to zero
4Scan all seven engines togetherCross-engine gaps are the finding, and they are invisible one engine at a time
5Record per prompt, not just the totalThe aggregate cannot tell you which page to write
6Re-run on a fixed cadenceMonthly for trend. A single scan is a snapshot, not a signal

Localise from the start if you sell in the UK. British phrasing pulls different sources than American phrasing for the same question, and a prompt set written in US English will quietly under-report the UK content you have already published.

What actually moves the score?

Two workstreams, on very different clocks. The mistake is sequencing them, because the slow one takes months and only starts when you start it.

On-page, measurable in weeks. Put a direct answer to the buyer's question in the first sentence under each heading, in server-rendered HTML. Many AI crawlers do not execute JavaScript, so a comparison table rendered client-side does not exist from the engine's point of view. Keep URLs short: in our data, pages with 1 to 3 word slugs earn 9.6 citations each at a 48.3% mention rate, against 5.6 citations and 19.5% for slugs of 10 words or more.

Off-page, measurable in months. This is where most of the score actually lives, and it is the part teams underinvest in. Across the 267,963 citations in our earlier off-page study, 84% pointed at pages the mentioned brand does not own.

84%
of citations point at pages the brand does not own
64.6%
of directory citations convert to a brand mention
5.2%
of blog citations convert to a brand mention
67x
more citations per URL on our own homepage than our own blog

That directory figure is the one that changes plans. Sorted by how reliably a citation turns into an engine actually naming you:

Source typeCitationsMention rate
Reference and wiki-like7,04765.1%
Directory19,06864.6%
Review site4,21162.8%
Marketplace8,78057.7%
Your own website229,65136.6%
Reddit2,79437.7%
Blog1,6195.2%

Directories, review sites and marketplaces convert citations into mentions at roughly 1.8 times the rate of your own website and 12 times the rate of a blog post. A G2 or Capterra listing does more for a visibility score than several months of blog output, and it is a one-week job.

This is not an argument against publishing. It is an argument against expecting the blog alone to move a composite whose largest component is off-page.

How long should each change take to show up?

Setting the expectation correctly is what stops a programme being cancelled in month two.

ChangeRealistic lagWhat you should see first
Fixing a blocked crawlerDays to 2 weeksCitations appear where there were none
Adding direct answers under headings2 to 4 weeksProminence improves before mention rate does
Shortening URLs on new pagesNext crawl cycleApplies to new pages, not retrospectively
A new directory or review listing4 to 8 weeksMention rate moves on comparison prompts
Earned editorial coverage3 to 6 monthsShare of voice, the slowest and most durable component

Measure each against the specific prompts the change was meant to affect, not the headline score. Cluster-level movement is attributable. Headline movement is too noisy to tell you anything about a single page.

Key takeaways

PointDetails
The score is a smoke alarmIt tells you whether you have a problem. The per-prompt detail underneath tells you which page to write.
Four components, four different fixesMention rate, prominence, citation rate and share of voice come apart. Diagnose which is low before acting.
Engines disagree structurallyChatGPT takes 34.1% of citations from homepages and 5.2% from blogs. Google's surfaces are close to the reverse.
Never compare across toolsA 15 to 60 point spread between vendors on the same brand is normal. Compare a score only to its own history.
Off-page carries most of the weight84% of citations point at pages you do not own. Directories convert at 64.6%, blogs at 5.2%.
Run both clocks at onceOn-page shows in weeks, earned coverage in months. Sequencing them wastes the slow one.

A practitioner's view on what the number is worth

I spent a decade signing off marketing budgets before I worked on this, so my first question about any new metric is what decision it changes. For an AI visibility score the honest answer is that it changes a lot in month one and very little by month four. In month one it is genuinely diagnostic: you find out whether engines know you exist, and that reframes the plan. By month four you are watching a line move two points and building a story about why, which is not analysis, it is decoration.

So my rule is that the composite is a reporting artefact and the prompt cluster is the working unit. I almost never open the headline number first. I open the list of prompts where a competitor is named and we are not, because that list is a content brief that writes itself and the score is just a lagging summary of how well we worked through it.

The mistake I made early was treating a low citation rate as a content quality problem, which is the expensive way to read it. It is usually a distribution problem wearing a content costume. We funded page improvements that were not the constraint, when the real one was that almost nobody else on the internet was describing us, so engines had nothing to cite except our own claims about ourselves. The 84% figure is what changed my mind, and it should probably change yours: most of what an engine knows about your brand, it learned somewhere you do not control.

The uncomfortable part, if you are the person defending the budget, is that the work which moves the number is the work that reports worst. Publishing is visible and attributable. Getting listed, reviewed and written about by other people is slow, lumpy and hard to put in a monthly deck. It is still where the number moves, and pretending otherwise just means funding the wrong half for longer.

Where Koalr fits

Koalr runs 25 real buyer prompts from your category across all seven engines and returns an AI Visibility Score from 0 to 100, with the per-answer record behind it rather than just the number.

The parts that matter for the workflow above: Prompt Explorer shows exactly which questions name you and which name a rival, Sources ranks the domains and URLs the engines actually cited in your category, Competitors lines your score up against the brands appearing beside you, and GEO Audit scores how readable your pages are to a machine that does not run JavaScript. Actions turns each gap into a task rather than a chart.

Plans start at £79 a month with a 7-day free trial, and the free scan returns a baseline across all seven engines before you pay anything. If the scan shows a category where rivals own every answer, that is the finding, and it changes the plan more than any feature list will. Questions about your score go to hello@koalr.ai.

Frequently asked questions

There is no universal threshold, because the score is relative to your category and prompt set. A 30 in a crowded category where the leader holds 45 is a different situation from a 30 where the leader holds 90. Judge it against named competitors on identical prompts and against your own previous scans, never against an absolute benchmark.

Sources

Written by

Joe Creighton

Joe Creighton

Co-founder, Koalr

Joe is co-founder of Koalr. A decade as CFO and COO in enterprise software taught him to read a marketing report backwards: not what it claims, but where the number came from and what it quietly left out. He brings the same scrutiny to AI search, a channel most companies are currently running blind.

See how your brand shows up in AI search.

Enter your website and get your AI Visibility Score across the engines your buyers actually ask. First score in under 5 minutes, zero setup.

Koalr
© 2026 Koalr Ltd. All rights reserved.

Legal

Get started