Back to blog

AI Visibility

Where Does ChatGPT Get Its Information?

Three different places, and which one it uses changes the answer you get. Our citation data also shows ChatGPT sourcing the open web very differently from Google's AI surfaces, in a way that is visible in the numbers.

Jake ChurcherJake Churcher··8 min read
Where Does ChatGPT Get Its Information?
On this page

ChatGPT draws on three separate sources: the pretrained weights it learned during training, live web pages it fetches at the moment you ask, and licensed data feeds from publishers OpenAI has agreements with. Which one answers your question depends on what you asked and whether search was triggered. Only the second and third produce a citation you can see.

What are ChatGPT's three sources?

They behave differently enough that treating them as one thing is the root of most confusion about why answers change.

SourceWhat it holdsHow currentCited?
Pretrained weightsPatterns learned from a large web crawl during trainingFrozen at the training cutoffNo
Live retrievalPages fetched by the search tool while answeringCurrent to the minuteYes, as links
Licensed feedsContent OpenAI pays publishers forVaries by agreementSometimes

The practical consequence: if you ask about your company and get a confident answer with no links, you are reading the model's memory of a crawl that happened months ago. Correcting your website will not change that answer until either the model is retrained or the question starts triggering retrieval.

When does ChatGPT browse, and when does it answer from memory?

Retrieval fires on questions the model treats as time-sensitive, specific, or beyond what it confidently knows. Broad definitional questions usually answer from weights. Questions naming a company, a price, a comparison or a recent event usually trigger a search.

This matters more than it sounds, because the two paths reward completely different work. Retrieval rewards pages that exist, are crawlable and answer the question directly. Weights reward having been widely written about, everywhere, before the cutoff.

Most buyer questions in our tracked sets fall on the retrieval side, which is the more tractable half.

Which sites does ChatGPT actually cite?

Here the data gets interesting, and it is not what the popular studies suggest.

Across the 267,963 citations in our set, we looked at the five most-cited social and community domains and counted how many citations each engine contributed to them:

Citations to the top five social and community domains, by enginecitations
Google AI Mode4,450
Google AI Overviews4,156
Perplexity1,567
ChatGPT368
Copilot72
Claude3

ChatGPT contributes 368 citations to that group, against Google AI Mode's 4,450. On its own that reads as ChatGPT simply avoiding social sources. The breakdown is stranger than that: almost all of ChatGPT's 368 are Reddit. YouTube, Facebook and Instagram are close to zero.

That is not the shape of a model that dislikes social content. It is the shape of a model with access to one social platform and not the others, which is exactly what a content licensing agreement produces.

Claude, at 3 citations, has effectively opted out of citing these sources altogether.

Why isn't Wikipedia the top source?

Because most published research on this question measures encyclopedic queries, and buyers do not ask encyclopedic questions.

Wikipedia does not appear in the top 20 sources for the buyer-intent prompts in our data, which contradicts the widely-repeated finding that it dominates ChatGPT's sources. Both can be true. Ask "what is generative engine optimisation" and a reference source is a sensible thing to cite. Ask "which tool should I buy for this" and it is useless, so the engine reaches for round-ups, communities and vendor pages instead.

The lesson is to check which question a study measured before acting on it. We covered the full source ranking in which pages AI engines cite when they mention your brand.

How we counted, and why breadth beats volume

One methodology note, because it changes the answer.

If you rank sources by raw citation count, the results are dominated by whichever category happens to be largest in your sample. A single big vertical drags its favourite domains to the top and they look universal when they are not.

We rank by breadth instead: how many distinct brand categories a domain earns citations in. YouTube and Reddit appear in all 16 categories we measured. Google Maps appears in 15, Facebook in 14. A domain cited 3,000 times inside one category and nowhere else is a niche result, not a pattern, and ranking by volume would hide that distinction entirely.

Any source study that does not say how it ranked is worth reading sceptically.

See where your brand shows up in AI search

Enter your website and get your AI Visibility Score across 7 engines. First score in under 5 minutes, no setup.

What this means if you want to be cited

Four things follow directly from the above.

Assume retrieval, not memory. You cannot influence weights on any useful timescale. You can influence what a crawler finds today, so optimise for the path that is actually reachable.

Do not generalise from one engine. The bars above are the argument. Coverage across engines is not thoroughness for its own sake; it is the only way to know which retrieval set you are missing from.

Reddit is doing disproportionate work for ChatGPT specifically. Not as a place to post promotional content, which communities correctly punish, but as a place where genuine category discussion already happens and gets read.

A missing citation is not a missing mention. The engine can name you while citing somebody else's page, and that gap is where most of the addressable work sits.

FAQ

Frequently asked questions

No. Retrieval fires when the model treats a question as time-sensitive, specific, or beyond what it confidently knows. Broad definitional questions are usually answered from pretrained weights with no search and no citations. Questions naming a company, a price or a comparison are much more likely to trigger a live search.

Sources

Written by

Jake Churcher

Jake Churcher

Co-founder, Koalr

Jake is co-founder of Koalr, where he works on GEO and AXO: how brands get found, cited, and recommended by AI. A decade in IT productising and marketing services, now applied to AI search. Passionate about building AI tools that help businesses win.

See how your brand shows up in AI search.

Enter your website and get your AI Visibility Score across the engines your buyers actually ask. First score in under 5 minutes, zero setup.

Koalr
© 2026 Koalr Ltd. All rights reserved.

Legal

Get started