Back to blog

AI Visibility

Prompt Tracking: Fewer Questions, Watched Properly

Most prompt sets are too big to act on and too small to be representative. Here is how to choose the questions worth watching, how often to run them, and what to do the week a competitor appears and you do not.

Jake ChurcherJake Churcher··12 min read
Prompt Tracking: Fewer Questions, Watched Properly
On this page

Prompt tracking is the practice of running a fixed set of buyer questions through AI answer engines on a schedule, and recording whether your brand is named, where in the answer, with what sentiment, and which pages the engine cited. It is not keyword rank tracking with new labels, and it is not the prompt observability engineering teams use to watch token counts and latency. The unit of measurement is a question a customer would actually ask.

PointDetails
Track unbranded questionsBranded permutations tell you what you already know. The questions that decide new business are the category ones you are absent from.
A lean set beats a large oneTwenty prompts you act on beat two hundred you scroll past. Expand only after you have worked half the signals the current set produced.
One engine is not a sampleAcross 295,746 citations, ChatGPT takes 34.1% of its citations from homepages and 5.2% from blog paths. Google's surfaces do close to the opposite.
Run each prompt more than onceAnswers move between sessions. A single run is an anecdote; the modal result of three is a reading.
The output is a brief, not a chartA prompt where a rival is named and you are not is a content brief with the research already done.

What is prompt tracking, and what is it not?

Prompt tracking monitors how AI answer engines describe your brand for a saved set of questions. Four things get recorded per run: whether you are mentioned, how prominently, what sentiment attaches to the description, and which URLs the engine cited as its sources. Run the same set on a schedule and you get a trend rather than an anecdote.

The name collides with something else, and the collision is expensive. Engineering teams use prompt tracking to mean LLM observability: prompt versions, token counts, latency, cost per call. Datadog's LLM Observability product is the well-known example. Those metrics matter enormously if you are building an AI feature. They tell you nothing about whether a buyer asking Gemini for a recommendation is shown your brand.

Brand prompt trackingLLM prompt observability
Question it answersDo engines name us when buyers ask?Is our AI feature fast, cheap and stable?
UnitA buyer questionAn API call
MetricsMentions, prominence, citations, sentiment, share of voiceTokens, latency, cost, error rate
BuyerMarketingEngineering
OutputA content briefA performance fix

Both are legitimate. Buying the second when you needed the first is a common and quiet failure, because the dashboard looks impressive and produces nothing anyone in marketing can act on.

Where do good prompts come from?

From places where somebody has already used the words. Invented prompts test your vocabulary rather than your buyers', and they inflate your score because you unconsciously phrase them the way your own site is written.

Five sources, in the order they tend to pay off:

  1. Sales call recordings and support tickets. The highest-fidelity source there is, because the phrasing is verbatim and the intent is confirmed. Take the question as spoken, including the hedges and the run-on clauses.
  2. Search Console queries, filtered to non-brand. Real language, already attached to your category, with the branded noise stripped out.
  3. Your own lost-deal notes. The comparison a prospect made out loud before choosing someone else is the exact question you want an engine asked.
  4. People Also Ask boxes on your core category terms. Cheap, plentiful, and reflective of follow-on intent.
  5. Competitor gap probes. Ask the category question without naming anyone and record who the engine names. This is the source that produces briefs fastest.

Then prune, which matters more than sourcing. Three cuts:

  • Cut branded permutations. If the question contains your name, the engine has already been told the answer. Keep two or three for reputation monitoring and drop the rest.
  • Cut near-duplicates. If two prompts return substantially the same answer and cite the same pages, they are one prompt. Keep the phrasing closest to how a buyer actually speaks.
  • Cut awareness-stage questions you cannot serve. "What is a CRM" is a question you will never win and would not benefit from winning.

How many prompts, and how often?

Koalr runs 25 real prompts per scan, and that number is deliberate rather than arbitrary. It is large enough that one odd answer does not move the aggregate, and small enough that a marketing team can read every result in a sitting and act on the interesting ones.

Programme stagePrompt setCadenceWhat you are looking for
Baseline, weeks 1 to 415 to 25Daily or near-dailyHow much answers move on their own, before you change anything
Steady state25 to 50WeeklyTrend, and competitors gaining on specific questions
Multiple products or markets50 to 150, segmentedWeekly per segmentWhich segment is weakest, not which prompt
Agency portfolioA separate set per clientWeeklyClient-level reporting, never a pooled number

The daily start is the part teams skip and then regret. You cannot tell whether a change you made worked unless you know how much the number moves when you do nothing at all. Four weeks of daily runs on a small set gives you that baseline. After it, weekly is enough for most B2B categories.

The failure mode in the other direction is more common. A 200-prompt set tracked weekly generates more findings than any team can work, so nobody works any of them, and the programme quietly becomes a dashboard nobody opens.

49.6%
of one publisher's citations came from just five URLs, out of roughly 1,549 pagesKoalr platform data

Citations concentrate hard. The same asymmetry applies to prompts: a handful will carry most of your commercial exposure, and the rest are texture. Finding which handful is the entire point of the baseline period.

Which engines does the set need to cover?

All seven, because they disagree with each other more than most teams expect, and the disagreement is structural rather than random. Koalr covers ChatGPT, Claude, Gemini, Perplexity, Microsoft Copilot, Google AI Overviews and Google AI Mode.

Across 295,746 citations, the page type each engine prefers to cite splits cleanly:

Page typeChatGPTCopilotPerplexityGeminiAI ModeAI Overviews
Homepage34.1%30.0%10.9%5.6%8.2%9.3%
Product or feature page11.0%4.0%21.2%10.5%23.7%24.4%
Blog post5.2%12.1%17.3%12.8%24.0%26.1%
Question-style page7.8%3.9%19.8%7.2%27.2%29.8%

Read across the blog row. Google's two surfaces take half their citations from blog and question-style pages; ChatGPT takes 5.2% from blog paths. A team that tracks only ChatGPT and concludes its content programme is failing has measured the one engine that was never going to reward it. A team that tracks only AI Overviews reaches the opposite wrong answer.

This is why per-engine reporting matters more than a blended score. The blend tells you something is wrong. The split tells you what.

What does the report actually need to show?

Six columns, and prominence is the one teams leave out and then miss.

MetricWhat it saysAct when
Mention rateShare of prompts where you are named at allIt falls two periods running
ProminenceWhere you land in the answer, first or fifthYou slip out of the first recommendation
Citation rateWhether the engine attributes a page of yoursIt stays flat while mentions rise
SentimentHow the description frames youAnything negative on a purchase-stage prompt
Share of voiceYour mentions against everyone namedA rival gains on high-intent prompts specifically
Cited sourcesWhich pages the engine usedThe cited page is not yours, or is wrong about you

Mention rate rising while citation rate stays flat is the pattern worth understanding, because it looks like progress and often is not. It means engines are learning about you from someone else's page. In our data, 84% of the 267,963 citations we have recorded point at pages the mentioned brand does not own, so this is the normal state rather than an anomaly. The buyer reads your name and then clicks through to a page you do not control.

What does the workflow look like week to week?

  1. 1

    Freeze the prompt set and version it

    Record every wording change with a date and a reason. Without that log you cannot tell whether a visibility shift came from your content, from a model update, or from you quietly rephrasing the question.

  2. 2

    Run every prompt more than once

    Three runs per prompt per session, and take the modal answer. Single runs carry more variance than teams expect, and a false positive costs a wasted content brief.

  3. 3

    Record the answer, not just the verdict

    Store the raw answer text and the cited URLs alongside the scores. The scores tell you a gap exists; only the text tells you why the engine chose who it chose.

  4. 4

    Compare against the baseline, not against zero

    Movement inside the range you measured during the baseline period is noise. Movement outside it is a signal worth spending on.

  5. 5

    Convert gaps into briefs

    A prompt where a rival is named and you are not already contains the research: the question, the competing answer, and the page the engine trusted.

  6. 6

    Re-measure the prompt you targeted

    Not the aggregate. The specific prompt the work was aimed at, over two consecutive periods, with the content change dated in the same log.

How do you turn a gap into a content brief?

Three patterns come up repeatedly, and each has a different fix. Diagnosing which one you are looking at takes about a minute once you have the answer text in front of you.

You are absent entirely. No mention on three or more engines, and no page of yours addresses the question directly. This is a new-page brief. Specify the question the page answers in its first 40 to 60 words, the schema type, and at least one number only you can publish.

You are named but placed last. The engine knows you and ranks you low. That is rarely a content problem on your own site. It is usually that the third-party sources the engine trusts describe your competitors more confidently than they describe you. The fix is off-page, and it is slower.

You are cited but not named. The engine used your page and credited the URL without writing your brand into the answer text. Usually the page never states plainly who published it. Adding a clear first-sentence brand statement to the cited page often resolves this inside a few weeks.

Before you brief anything

  • Confirm the gap held across two consecutive runs, not one
  • Read the winning answer in full and note which source it leaned on
  • Check whether an existing page of yours already targets the question badly
  • Decide whether the fix is on-page, off-page, or entity clarity
  • Write the target prompt into the brief so the re-measurement is unambiguous

The off-page case deserves a note, because it is the one teams under-resource. Measured across our citation set, source types convert to brand mentions at wildly different rates: directories at 64.6%, review sites at 62.8%, marketplaces at 57.7%, and blog-type sources at 5.2%. If prompt tracking keeps showing you named-but-last, another blog post is unlikely to be the lever.

See where your brand shows up in AI search

Enter your website and get your AI Visibility Score across 7 engines. First score in under 5 minutes, no setup.

What should you expect it to cost, in effort?

Less than teams fear on setup, more than they expect on the reading.

ActivityRealistic effort
Sourcing and pruning the first 25 promptsHalf a day, once
Baseline periodAutomated, four weeks of elapsed time
Weekly review of results30 to 45 minutes
Monthly manual read of sample answersOne hour
Acting on a briefWhatever your normal content cycle costs

The weekly review is the line item that decides whether the programme works. Tooling makes the collection free and the interpretation is still yours. A team that automates tracking and never reads the answers has bought a very precise way of not knowing anything.

Koalr's plans start at £79 a month for Starter, £199 for Growth and £349 for Business, each with a 7-day free trial and all seven engines included on every tier.

A practitioner's view on prompt tracking trade-offs

The mistake I made first, and see most often, was treating coverage as the goal. I wanted the set to be representative, so it grew, and the larger it got the less anyone did with it. The corrective that worked was a rule rather than a number: do not add a prompt until you have acted on half the signals the current set has produced. That put a natural ceiling on the set at whatever size the team could actually work, which turned out to be far smaller than I expected.

The second trade-off is automation against judgement, and I think the honest answer is that automation is worse at the thing that matters most. It is excellent at volume, cadence and trend. It is poor at noticing that the answer naming you positively is recommending you for the wrong use case, or that the sentence describing you is technically accurate and commercially damaging. Those show up in the text, not the score, and the only way to catch them is to read a sample by hand every month. I have never regretted that hour.

The third is more uncomfortable. Prompt tracking measures a probabilistic system, and some of what you observe is model behaviour you cannot influence at all. A model update can move your numbers more than a quarter of content work. That is not a reason to skip the measurement, but it is a reason to be careful about attributing every rise to your own cleverness. Log the model updates alongside your content changes, and be willing to say a movement is unexplained.

Where Koalr fits

Koalr runs 25 real buyer prompts per scan across all seven engines, records mentions, prominence, sentiment and cited sources for each, and tracks the four scores over time so you can tell a trend from a bad day. Gaps arrive as a ranked list rather than a spreadsheet, with the competing answer and the source it leaned on attached.

The parts of this article that are manual stay manual on purpose. Sourcing prompts from your own sales calls, and reading a sample of the answer text each month, are judgement work. Everything around them is collection, and collection is what the tool is for.

FAQ

Frequently asked questions

Start with 15 to 25 and only expand once you have acted on at least half the signals that set produces. Koalr runs 25 per scan as its standard. A larger set is not more rigorous if nobody works the findings, and in practice a bloated set is the most common reason a programme dies.

Sources

  • Koalr platform data, 295,746 citations across 23 projects and six engines: the per-engine page-type table and the citation concentration figure. Re-derived quarterly from our own citation records.
  • Koalr platform data, 267,963 citations: the 84% off-page share and the source-type mention conversion rates.
  • Datadog LLM Observability: the reference point for engineering-side prompt observability, useful for seeing exactly how different its metric set is from a brand visibility report.
  • Which pages do AI engines cite when they mention your brand?: covers reading the cited-source column, which is the half of a prompt report most teams skip.
  • How to get cited by AI engines through content you don't own: the off-page play for the named-but-last pattern above.
  • AI Visibility Score: what it measures and how to move it: how the four metrics in the report table combine into a single number, and when that number misleads.

Written by

Jake Churcher

Jake Churcher

Co-founder, Koalr

Jake is co-founder of Koalr, where he works on GEO and AXO: how brands get found, cited, and recommended by AI. A decade in IT productising and marketing services, now applied to AI search. Passionate about building AI tools that help businesses win.

See how your brand shows up in AI search.

Enter your website and get your AI Visibility Score across the engines your buyers actually ask. First score in under 5 minutes, zero setup.

Koalr
© 2026 Koalr Ltd. All rights reserved.

Legal

Get started