Prompt tracking is the practice of running a fixed set of buyer questions through AI answer engines on a schedule, and recording whether your brand is named, where in the answer, with what sentiment, and which pages the engine cited. It is not keyword rank tracking with new labels, and it is not the prompt observability engineering teams use to watch token counts and latency. The unit of measurement is a question a customer would actually ask.
| Point | Details |
|---|---|
| Track unbranded questions | Branded permutations tell you what you already know. The questions that decide new business are the category ones you are absent from. |
| A lean set beats a large one | Twenty prompts you act on beat two hundred you scroll past. Expand only after you have worked half the signals the current set produced. |
| One engine is not a sample | Across 295,746 citations, ChatGPT takes 34.1% of its citations from homepages and 5.2% from blog paths. Google's surfaces do close to the opposite. |
| Run each prompt more than once | Answers move between sessions. A single run is an anecdote; the modal result of three is a reading. |
| The output is a brief, not a chart | A prompt where a rival is named and you are not is a content brief with the research already done. |
What is prompt tracking, and what is it not?
Prompt tracking monitors how AI answer engines describe your brand for a saved set of questions. Four things get recorded per run: whether you are mentioned, how prominently, what sentiment attaches to the description, and which URLs the engine cited as its sources. Run the same set on a schedule and you get a trend rather than an anecdote.
The name collides with something else, and the collision is expensive. Engineering teams use prompt tracking to mean LLM observability: prompt versions, token counts, latency, cost per call. Datadog's LLM Observability product is the well-known example. Those metrics matter enormously if you are building an AI feature. They tell you nothing about whether a buyer asking Gemini for a recommendation is shown your brand.
| Brand prompt tracking | LLM prompt observability | |
|---|---|---|
| Question it answers | Do engines name us when buyers ask? | Is our AI feature fast, cheap and stable? |
| Unit | A buyer question | An API call |
| Metrics | Mentions, prominence, citations, sentiment, share of voice | Tokens, latency, cost, error rate |
| Buyer | Marketing | Engineering |
| Output | A content brief | A performance fix |
Both are legitimate. Buying the second when you needed the first is a common and quiet failure, because the dashboard looks impressive and produces nothing anyone in marketing can act on.
Where do good prompts come from?
From places where somebody has already used the words. Invented prompts test your vocabulary rather than your buyers', and they inflate your score because you unconsciously phrase them the way your own site is written.
Five sources, in the order they tend to pay off:
- Sales call recordings and support tickets. The highest-fidelity source there is, because the phrasing is verbatim and the intent is confirmed. Take the question as spoken, including the hedges and the run-on clauses.
- Search Console queries, filtered to non-brand. Real language, already attached to your category, with the branded noise stripped out.
- Your own lost-deal notes. The comparison a prospect made out loud before choosing someone else is the exact question you want an engine asked.
- People Also Ask boxes on your core category terms. Cheap, plentiful, and reflective of follow-on intent.
- Competitor gap probes. Ask the category question without naming anyone and record who the engine names. This is the source that produces briefs fastest.
Then prune, which matters more than sourcing. Three cuts:
- Cut branded permutations. If the question contains your name, the engine has already been told the answer. Keep two or three for reputation monitoring and drop the rest.
- Cut near-duplicates. If two prompts return substantially the same answer and cite the same pages, they are one prompt. Keep the phrasing closest to how a buyer actually speaks.
- Cut awareness-stage questions you cannot serve. "What is a CRM" is a question you will never win and would not benefit from winning.
How many prompts, and how often?
Koalr runs 25 real prompts per scan, and that number is deliberate rather than arbitrary. It is large enough that one odd answer does not move the aggregate, and small enough that a marketing team can read every result in a sitting and act on the interesting ones.
| Programme stage | Prompt set | Cadence | What you are looking for |
|---|---|---|---|
| Baseline, weeks 1 to 4 | 15 to 25 | Daily or near-daily | How much answers move on their own, before you change anything |
| Steady state | 25 to 50 | Weekly | Trend, and competitors gaining on specific questions |
| Multiple products or markets | 50 to 150, segmented | Weekly per segment | Which segment is weakest, not which prompt |
| Agency portfolio | A separate set per client | Weekly | Client-level reporting, never a pooled number |
The daily start is the part teams skip and then regret. You cannot tell whether a change you made worked unless you know how much the number moves when you do nothing at all. Four weeks of daily runs on a small set gives you that baseline. After it, weekly is enough for most B2B categories.
The failure mode in the other direction is more common. A 200-prompt set tracked weekly generates more findings than any team can work, so nobody works any of them, and the programme quietly becomes a dashboard nobody opens.
Citations concentrate hard. The same asymmetry applies to prompts: a handful will carry most of your commercial exposure, and the rest are texture. Finding which handful is the entire point of the baseline period.
Which engines does the set need to cover?
All seven, because they disagree with each other more than most teams expect, and the disagreement is structural rather than random. Koalr covers ChatGPT, Claude, Gemini, Perplexity, Microsoft Copilot, Google AI Overviews and Google AI Mode.
Across 295,746 citations, the page type each engine prefers to cite splits cleanly:
| Page type | ChatGPT | Copilot | Perplexity | Gemini | AI Mode | AI Overviews |
|---|---|---|---|---|---|---|
| Homepage | 34.1% | 30.0% | 10.9% | 5.6% | 8.2% | 9.3% |
| Product or feature page | 11.0% | 4.0% | 21.2% | 10.5% | 23.7% | 24.4% |
| Blog post | 5.2% | 12.1% | 17.3% | 12.8% | 24.0% | 26.1% |
| Question-style page | 7.8% | 3.9% | 19.8% | 7.2% | 27.2% | 29.8% |
Read across the blog row. Google's two surfaces take half their citations from blog and question-style pages; ChatGPT takes 5.2% from blog paths. A team that tracks only ChatGPT and concludes its content programme is failing has measured the one engine that was never going to reward it. A team that tracks only AI Overviews reaches the opposite wrong answer.
This is why per-engine reporting matters more than a blended score. The blend tells you something is wrong. The split tells you what.
What does the report actually need to show?
Six columns, and prominence is the one teams leave out and then miss.
| Metric | What it says | Act when |
|---|---|---|
| Mention rate | Share of prompts where you are named at all | It falls two periods running |
| Prominence | Where you land in the answer, first or fifth | You slip out of the first recommendation |
| Citation rate | Whether the engine attributes a page of yours | It stays flat while mentions rise |
| Sentiment | How the description frames you | Anything negative on a purchase-stage prompt |
| Share of voice | Your mentions against everyone named | A rival gains on high-intent prompts specifically |
| Cited sources | Which pages the engine used | The cited page is not yours, or is wrong about you |
Mention rate rising while citation rate stays flat is the pattern worth understanding, because it looks like progress and often is not. It means engines are learning about you from someone else's page. In our data, 84% of the 267,963 citations we have recorded point at pages the mentioned brand does not own, so this is the normal state rather than an anomaly. The buyer reads your name and then clicks through to a page you do not control.
What does the workflow look like week to week?
- 1
Freeze the prompt set and version it
Record every wording change with a date and a reason. Without that log you cannot tell whether a visibility shift came from your content, from a model update, or from you quietly rephrasing the question.
- 2
Run every prompt more than once
Three runs per prompt per session, and take the modal answer. Single runs carry more variance than teams expect, and a false positive costs a wasted content brief.
- 3
Record the answer, not just the verdict
Store the raw answer text and the cited URLs alongside the scores. The scores tell you a gap exists; only the text tells you why the engine chose who it chose.
- 4
Compare against the baseline, not against zero
Movement inside the range you measured during the baseline period is noise. Movement outside it is a signal worth spending on.
- 5
Convert gaps into briefs
A prompt where a rival is named and you are not already contains the research: the question, the competing answer, and the page the engine trusted.
- 6
Re-measure the prompt you targeted
Not the aggregate. The specific prompt the work was aimed at, over two consecutive periods, with the content change dated in the same log.
How do you turn a gap into a content brief?
Three patterns come up repeatedly, and each has a different fix. Diagnosing which one you are looking at takes about a minute once you have the answer text in front of you.
You are absent entirely. No mention on three or more engines, and no page of yours addresses the question directly. This is a new-page brief. Specify the question the page answers in its first 40 to 60 words, the schema type, and at least one number only you can publish.
You are named but placed last. The engine knows you and ranks you low. That is rarely a content problem on your own site. It is usually that the third-party sources the engine trusts describe your competitors more confidently than they describe you. The fix is off-page, and it is slower.
You are cited but not named. The engine used your page and credited the URL without writing your brand into the answer text. Usually the page never states plainly who published it. Adding a clear first-sentence brand statement to the cited page often resolves this inside a few weeks.
Before you brief anything
- Confirm the gap held across two consecutive runs, not one
- Read the winning answer in full and note which source it leaned on
- Check whether an existing page of yours already targets the question badly
- Decide whether the fix is on-page, off-page, or entity clarity
- Write the target prompt into the brief so the re-measurement is unambiguous
The off-page case deserves a note, because it is the one teams under-resource. Measured across our citation set, source types convert to brand mentions at wildly different rates: directories at 64.6%, review sites at 62.8%, marketplaces at 57.7%, and blog-type sources at 5.2%. If prompt tracking keeps showing you named-but-last, another blog post is unlikely to be the lever.
See where your brand shows up in AI search
Enter your website and get your AI Visibility Score across 7 engines. First score in under 5 minutes, no setup.
What should you expect it to cost, in effort?
Less than teams fear on setup, more than they expect on the reading.
| Activity | Realistic effort |
|---|---|
| Sourcing and pruning the first 25 prompts | Half a day, once |
| Baseline period | Automated, four weeks of elapsed time |
| Weekly review of results | 30 to 45 minutes |
| Monthly manual read of sample answers | One hour |
| Acting on a brief | Whatever your normal content cycle costs |
The weekly review is the line item that decides whether the programme works. Tooling makes the collection free and the interpretation is still yours. A team that automates tracking and never reads the answers has bought a very precise way of not knowing anything.
Koalr's plans start at £79 a month for Starter, £199 for Growth and £349 for Business, each with a 7-day free trial and all seven engines included on every tier.
A practitioner's view on prompt tracking trade-offs
The mistake I made first, and see most often, was treating coverage as the goal. I wanted the set to be representative, so it grew, and the larger it got the less anyone did with it. The corrective that worked was a rule rather than a number: do not add a prompt until you have acted on half the signals the current set has produced. That put a natural ceiling on the set at whatever size the team could actually work, which turned out to be far smaller than I expected.
The second trade-off is automation against judgement, and I think the honest answer is that automation is worse at the thing that matters most. It is excellent at volume, cadence and trend. It is poor at noticing that the answer naming you positively is recommending you for the wrong use case, or that the sentence describing you is technically accurate and commercially damaging. Those show up in the text, not the score, and the only way to catch them is to read a sample by hand every month. I have never regretted that hour.
The third is more uncomfortable. Prompt tracking measures a probabilistic system, and some of what you observe is model behaviour you cannot influence at all. A model update can move your numbers more than a quarter of content work. That is not a reason to skip the measurement, but it is a reason to be careful about attributing every rise to your own cleverness. Log the model updates alongside your content changes, and be willing to say a movement is unexplained.
Where Koalr fits
Koalr runs 25 real buyer prompts per scan across all seven engines, records mentions, prominence, sentiment and cited sources for each, and tracks the four scores over time so you can tell a trend from a bad day. Gaps arrive as a ranked list rather than a spreadsheet, with the competing answer and the source it leaned on attached.
The parts of this article that are manual stay manual on purpose. Sourcing prompts from your own sales calls, and reading a sample of the answer text each month, are judgement work. Everything around them is collection, and collection is what the tool is for.
FAQ
Frequently asked questions
Start with 15 to 25 and only expand once you have acted on at least half the signals that set produces. Koalr runs 25 per scan as its standard. A larger set is not more rigorous if nobody works the findings, and in practice a bloated set is the most common reason a programme dies.
Sources
- Koalr platform data, 295,746 citations across 23 projects and six engines: the per-engine page-type table and the citation concentration figure. Re-derived quarterly from our own citation records.
- Koalr platform data, 267,963 citations: the 84% off-page share and the source-type mention conversion rates.
- Datadog LLM Observability: the reference point for engineering-side prompt observability, useful for seeing exactly how different its metric set is from a brand visibility report.
- Which pages do AI engines cite when they mention your brand?: covers reading the cited-source column, which is the half of a prompt report most teams skip.
- How to get cited by AI engines through content you don't own: the off-page play for the named-but-last pattern above.
- AI Visibility Score: what it measures and how to move it: how the four metrics in the report table combine into a single number, and when that number misleads.
