An AI visibility tool has to prove three things before it earns a contract: it runs the questions your buyers actually type, it checks every engine those buyers use, and it turns what it finds into something you can fix. Everything past that is presentation on top of those three jobs, and most demos spend forty minutes on the presentation.
There is no test drive for most of these tools before you sign. That is what makes the questions do the work a trial normally would.
Buyer asks: “what should I check before I buy an AI visibility tool”
Ask three things before you sign anything. First, request the exact prompt list the tool will run for your category, not a sample: a canned list built for a demo tells you nothing about your buyers. Second, confirm each AI platform is checked and reported on separately, since "Google" can mean three different products with three different answers. Third, ask what happens after the tool finds a gap: a score with no next step is trivia, not a plan.
Generic, vendor-agnostic advice. No brand named, on either side.
What does an AI visibility tool actually need to do?
It has to run real buyer questions against every engine that matters to you, on a repeatable schedule, and hand back something more useful than a badge. A tool that nails the first two and stops at a dashboard number has built half a product. The other half, turning a gap into a page you can write or fix, is where most of the category's actual value sits.
Five things separate a tool worth paying for from one that looks the part in a sales call:
| Criterion | Why it matters | What to check |
|---|---|---|
| Real buyer prompts | A tool built on guessed keywords answers a question nobody asked | Ask to see the exact prompt list before you sign, not a sample deck |
| Every engine, named separately | ChatGPT, Gemini and Google AI Overviews cite different pages for the same question | Check whether "Google" on the feature list means one surface or three |
| Citation and mention, as two numbers | A page can be quoted as a source without your brand ever being said aloud | Ask whether the tool reports both, or collapses them into one score |
| A fix, not just a number | A gap you cannot act on is trivia | Look for a prioritised list of what to do next, not only a 0-100 badge |
| Multi-brand handling | Agencies price per client differently to a single in-house brand | Ask exactly what a second brand or client costs, in writing |
Real prompts beat guessed ones
A prompt list built from keyword-tool volume misses most of what buyers actually ask. Real questions run long, specific and conversational: "which tool tracks where my website gets cited by ai answer engines" rather than "ai visibility tool." Keyword tools grade the second phrasing and miss the first, because search-volume estimators are built for typed queries, not for the way people talk to a chat window.
Ask a vendor to show you the literal prompts it will run for your category before you commit. If they can only describe the process in the abstract, generic questions, standard categories, industry benchmarks, you are being sold a workflow, not a measurement built for your buyers.
Which engines does it actually check?
7 platforms answer buyer questions today: ChatGPT, Claude, Gemini, Perplexity, Microsoft Copilot, Google AI Overviews and Google AI Mode. A tool that lists "Google" as 1 line item is bundling 3 separate products that behave nothing alike.
Gemini is the standalone chat app. Google AI Overviews is the summary panel above ordinary search results. Google AI Mode is the conversational tab inside Search. A brand can be named clearly in one of the 3 and absent from the other two, because each pulls from a different mix of sources to build its answer. Our breakdown of tracking a brand across ChatGPT and Gemini covers why that split matters in practice, and how AI Overviews specifically decide what to cite is worth reading before you assume "Google coverage" means what the pricing page implies.
Ask the vendor, in writing, whether their reporting breaks Google into three lines or one. It is the fastest way to find out whether a feature list was written by someone who has actually used the product.
Citation and mention are not the same number
A citation is an engine using your page as a source. A mention is the engine saying your brand's name in the answer it gives the reader. The two move independently: a page can be pulled in as a reference and never once get the brand named aloud, and a brand can be recommended by name from a source that never linked back.
A tool that reports a single blended "visibility score" is hiding which of those two things actually happened. If you were cited but not mentioned, the fix is usually about how your page is written. If you were mentioned but never cited, the story is closer to reputation than content. Those are different jobs for different teams, and you cannot tell them apart from one number.
See where your brand shows up in AI search
Enter your website and get your AI Visibility Score across 7 engines. First score in under 5 minutes, no setup.
A score without a fix is trivia
The number is the least useful part of the report. What makes an AI Visibility Score worth checking weekly is the list underneath it: which questions a competitor is winning, which page they are winning it with, and what that implies you should write or change.
That is the difference between a tracker and a tool that pays for itself. Koalr's scoring model exists to feed a prioritised action list, not to be the headline number on a slide. When you evaluate a vendor, ask them to show you what happens the day after a gap shows up on the dashboard. If the honest answer is "you'll know," keep looking.
If you run more than one brand, check the maths first
Agencies and portfolio teams price this category very differently to a single in-house brand, and the difference is easy to miss until the invoice arrives. Some vendors charge full price per client with no shared pool of prompts. Others cap the engines available below a top-tier plan, so the client roster you actually run includes at least one brand missing an engine a buyer might ask.
Before you commit a client roster to one vendor, get the per-client cost in writing, not the headline plan price, and confirm every engine on the list is included at the tier you are actually buying. Our comparison of AI visibility tools built for agencies walks through what five vendors' own pricing pages say once you add the engines a real roster needs.
The question I would ask before signing anything
Ask for the prompt list before the demo, not after. If a vendor cannot show you the exact questions it will run for your category, in your market's language, you are looking at a workflow, not a measurement.
Here is the part that is inconvenient to say and true anyway: no tool, this one included, can promise an engine changes its answer. Nothing sits between you and the model's own retrieval and ranking. What a good tool can promise is showing you precisely which page a competitor is winning with on a question you should own, so you spend the next sprint fixing the right thing instead of guessing at eight things that might matter.
That is a narrower promise than most sales pages make. It is also the only one a vendor can actually keep.
Common mistakes buyers make
Three show up again and again in how teams pick this category short.
Judging on the score alone, before checking what sits underneath it. A polished 0-100 number with no prioritised list attached is a headline, not a plan.
Assuming an existing SEO or social-listening tool already covers this. Rank trackers watch Google's ten blue links. Listening tools watch what humans say about you online. Neither one reads what an AI engine says back when a buyer asks it a direct question, which is a different surface with different sources.
Signing before seeing the real prompt list. A demo built on generic category questions proves the interface works. It proves nothing about whether the tool will catch the specific way your buyers ask.
Where to check your own answer
The fastest way to test any of this is against your own brand. Run a free scan at koalr.ai: it takes a domain, no card, and returns a first score across all seven engines in under five minutes. Questions about what the report shows go to hello@koalr.ai.
Frequently asked questions
Software that runs real buyer questions against AI answer engines like ChatGPT, Gemini and Perplexity on a schedule, then reports whether your brand is cited as a source, named in the answer, or missing entirely. It answers a different question to an SEO rank tracker: not where your pages rank, but what the engine actually says when a buyer asks.

