CITERA :: ARTICLE
Where AI engines actually source their answers
How ChatGPT, Perplexity, AI Overviews and Gemini each assemble a response, and which source types keep appearing across categories.
Author
Liam Karlsson
In this article
The two ways a model knows about you
Engine by engine
The source types that keep appearing
What this changes about the work
How to verify any of this yourself
Every AI answer is built from something. Understanding what that something is — per engine — explains most of what looks arbitrary about AI visibility.
This piece breaks down how the major answer engines assemble a response, and what that means for where your brand needs to exist.
The two ways a model knows about you
There are only two, and they behave completely differently.
Parametric memory is what the model absorbed during training. It is frozen at the training cutoff, has no citations, and cannot be updated except by retraining. If a model “knows” your brand this way, that knowledge may be years stale and there is no page you can edit to fix it.
Retrieval is what the engine fetches at question time — live search results, an internal index, or a partner’s index. This is where nearly all practical AEO work happens, because retrieval reads pages as they are today.
Most answers blend the two. A model recalls a category from memory, then retrieves current pages to fill in specifics. This is why a brand can be named confidently but described with outdated pricing.
Engine by engine
ChatGPT
Browsing-enabled responses retrieve live results and cite sources inline. Without browsing, answers come from parametric memory alone — no citations, and a knowledge cutoff that can be well behind the present. Practically, this means the same question can produce a cited, current answer or an uncited, stale one depending on whether the engine decided a search was warranted. [1]
What this rewards: being present in mainstream indexed sources that a search step would surface, and being described consistently enough that the memorised version and the retrieved version agree.
Perplexity
Retrieval-first by design. Perplexity runs searches, reads the returned pages, and answers with dense inline citations. Answers are typically grounded in a handful of pages per response, and those pages are visible to the user. [2]
What this rewards: pages that directly and completely answer a specific question. Perplexity is the engine where an obscure but precise page most often beats a famous but vague one.
Google AI Overviews
Grounded in Google’s own index and shaped by what already ranks. The pages cited in an Overview are usually drawn from results that were already competitive for the query.
What this rewards: conventional SEO strength, plus content structured so a specific passage can be lifted cleanly. If you rank nowhere for a query, you are unlikely to appear in its Overview.
Gemini
Grounded in Google Search results, with the model’s own knowledge layered over the top. Behaviour sits between ChatGPT and AI Overviews — more willing than Overviews to synthesise beyond the cited pages, more grounded than an unbrowsed ChatGPT response.
The source types that keep appearing
Across engines and categories, the same kinds of pages do a disproportionate amount of the work:
Community discussion — Reddit, Stack Overflow, niche forums. Heavily used for opinion-shaped questions: “is X any good”, “what do people actually use”
Review and directory sites — G2, Capterra, Trustpilot and vertical equivalents. Used for comparison and shortlist questions
Reference sites — Wikipedia and similar. Used for definitional and entity questions, and disproportionately trusted
Documentation — official docs rank well for how-to and capability questions, often above marketing pages from the same company
Independent editorial — trade publications and analyst-style writeups, used for category framing
Vendor sites — your own pages, usually cited for specifics like pricing and features rather than for claims about relative quality
The pattern worth internalising: engines cite you for facts about you, and cite others for judgements about you. You cannot self-declare your way into being called the best option.
What this changes about the work
If retrieval decides most answers, then AEO is less about persuading a model and more about making sure the right pages exist and say the right thing.
Three implications:
Third-party presence is not optional. The comparison and shortlist questions — the commercially valuable ones — are answered largely from sources you do not own.
Specificity beats volume. A page that completely answers one narrow question is retrievable. Fifty pages that partially answer it are not.
Stale facts are actively dangerous. A wrong price in an old listing does not just fail to help; it gets repeated as fact by a system users trust.
How to verify any of this yourself
Ask any engine a question in your category and read the citations rather than the answer. Do it across ten questions and the source pattern for your category becomes obvious within an hour. It will differ from ours, and yours is the one that matters.
Summarize this article with AI
Hand this page to your assistant of choice — the prompt is pre-filled:
ChatGPT · Claude · Perplexity · Gemini · Copilot · Grok · DeepSeek · Mistral · Qwen