Blog/AI Engines
AI Engines

How ChatGPT decides which brands to recommend

When ChatGPT names two products and ignores the rest of your category, that shortlist wasn't retrieved from a league table — it was assembled, on the fly, by four interacting mechanisms. Each one is a lever you can actually pull.

ML Maya Lindqvist · Head of Research June 30, 2026 12 min read
Key takeaways
  • ChatGPT's recommendations blend parametric knowledge (brand associations baked in at training time) with live retrieval — ChatGPT search rides a web index crawled by OpenAI's OAI-SearchBot.2
  • Consensus wins: a claim repeated consistently across independent sources gets treated as fact; a claim that lives only on your own domain gets treated as marketing.
  • Phrasing reshuffles the shortlist. "Best CRM" and "best CRM for a 5-person agency" activate different associations and retrieve different sources — you can rank in one and vanish from the other.
  • Answers are stochastic: the same prompt yields different brand sets run to run, so any serious measurement means repeated sampling, not one screenshot.

Ask ChatGPT to recommend a CRM, a CMS, a standing desk or an ERP consultant and you'll get a confident, tidy shortlist of two to four names. It reads like the output of some ranking system. It isn't. There is no brand league table inside the model — there is a probability distribution over words, a retrieval pipeline bolted to its side, and a sampling process that stitches them into prose.

If you want your brand in that prose, you need to understand the machine at the level where you can intervene. There are four mechanisms that matter: what the model remembers, what it looks up, how it weighs agreement between sources, and how randomness shuffles the result. Let's take them in order.

Mechanism 1: parametric knowledge — what the model remembers

Large language models are trained on enormous swaths of the public web. The GPT-3 paper documented the recipe openly: filtered Common Crawl, curated web text, books, and Wikipedia, with the curated and encyclopedic sources deliberately up-weighted relative to their raw size.1 Successor models have not published full corpora, but the shape is the same: the open web, filtered for quality, with heavier weight on sources the pipeline trusts.

During training, the model compresses all of that into weights — and your brand comes along for the ride. If documentation, reviews, comparison posts and forum threads repeatedly described your product as "the affordable option for small teams," that association is now part of the model's parameters. Ask a question that activates the "affordable + small team" region of its learned space, and your name becomes a high-probability next token.

Three properties of parametric knowledge matter strategically:

Intervention point: be well-described where models train. That means accurate, consistent brand descriptions across Wikipedia (where warranted), major review platforms, industry publications and community threads — not just your own site. This is the slow loop; treat it like brand PR with a two-model-generation payoff horizon.

Mechanism 2: retrieval — what ChatGPT looks up right now

Since the launch of SearchGPT and its merge into ChatGPT search, a large share of commercial queries no longer rely on memory alone. The model issues search queries against an index, fetches candidate pages, and synthesizes an answer with citations.

OpenAI operates distinct crawlers for distinct jobs, and the difference matters for your robots.txt:2

CrawlerWhat it feedsIf you block it
GPTBotTraining corpora for future modelsFuture models know less about you (parametric loop weakens)
OAI-SearchBotThe search index behind ChatGPT searchYou stop appearing as a linked, citable source in searched answers
ChatGPT-UserReal-time fetches when a user's request triggers a visitLive page reads on behalf of users fail

The retrieval loop is fast — publish a genuinely useful comparison page this week and it can be cited in answers next week. It is also where classic search discipline re-enters: if your page can't be crawled, parsed and excerpted, it can't be retrieved into the answer. Server-rendered HTML with the key claims in plain text beats a JavaScript app every time.

Intervention point: allow OAI-SearchBot explicitly, keep money pages fast and extractable, and lead each page with the direct, self-contained answer to the question the page exists for. The KDD 2024 GEO study quantified what retrieval rewards: adding citations, quotations and statistics lifted source visibility in generative answers by up to 40% across ~10,000 test queries.4

Mechanism 3: consensus — agreement beats assertion

Here is the mechanism most marketing teams underestimate. Both loops — training and retrieval — reward the same underlying signal: independent agreement.

In training, a fact that appears in one document is noise; a fact that recurs across thousands of documents becomes a stable statistical pattern in the weights. In retrieval, the model typically reads several fetched sources at once and synthesizes the points they agree on; a claim asserted by one source and contradicted (or simply unmentioned) by the others rarely survives into the final answer.

Practically, this means your own website has a ceiling. It's necessary — it's the canonical source for your specs and pricing — but a claim that exists only there reads to the machine like what it is: self-description. The same claim echoed in similar words by a review site, a comparison post, a Reddit thread and an industry publication reads like a fact about the world.

💡

Rule of thumb: you can't tell an LLM what you are. You can only arrange for many independent sources to describe you the same way — and then make sure your own pages agree with them, in quotable sentences.

Intervention point: run a description-consistency audit. Collect how the top ten third-party sources describe your product — category label, key differentiator, price point, target customer. Where they disagree with each other or with your site, fix the discrepancy at the source: brief analysts and reviewers with the same one-sentence positioning, update outdated listings, and use consistent phrasing in your own materials so there's a canonical sentence for others to echo.

See the consensus about you
What does ChatGPT actually say when buyers ask about your category?

MentionBeat runs your real buyer prompts against ChatGPT, Claude, Gemini and Perplexity, shows which brands the answers name, and traces the sources shaping them — so you know exactly where the consensus needs work.

Get a free visibility report

Mechanism 4a: phrasing — small wording changes, different shortlists

Ask "what's the best CRM?" and "what's the best CRM for a 5-person design agency?" and you are not asking the same machine question. The added qualifiers change which learned associations activate and — in searched answers — which queries get issued and which pages come back. A brand dominant in generic roundups can vanish from the qualified variant, and vice versa.

This is why single-prompt testing misleads. Real buyers phrase the same need dozens of ways: by budget ("cheapest…"), by role ("…for a marketing team"), by incumbent ("alternatives to X"), by constraint ("…that's HIPAA compliant"). Each phrasing family has its own winners.

Intervention point: build your prompt suite from phrasing families, not single keywords — and build content for the qualified variants you should own. A page that directly answers "best warehouse scanner for cold storage" can win that entire phrasing family precisely because nobody else bothered to answer it specifically.

Mechanism 4b: stochasticity — same question, different answers

Even with identical phrasing, ChatGPT samples its output token by token from a probability distribution. Run one commercial prompt ten times and the brand set shifts: the category leader shows up nearly every run, mid-consensus brands flicker in and out, and long-shot brands surface occasionally.

Brand A (leader)9/10
Brand B7/10
Brand C4/10
Your brand3/10
Brand D1/10

Illustrative example — not measured data. The pattern shows how one buying prompt, run 10 times against the same engine, typically yields a stable leader, flickering mid-tier brands, and occasional long-shots. Your real numbers will differ; the variance won't.

Two consequences follow. First, anecdotes are worthless: the screenshot where you appear and the screenshot where you don't are both "true," and neither is a measurement. What's real is the underlying probability — your mention rate — and you can only estimate it by sampling repeatedly and attaching confidence intervals. This is exactly the kind of repeated, scheduled sampling a platform like MentionBeat exists to automate.

Second, stochasticity is where challengers live. You rarely go from absent to ever-present in one step. You go from 0/10 to 3/10 to 6/10 as your consensus footprint and retrievability improve. A rising mention rate across samples is the earliest reliable signal that your GEO work is compounding — visible weeks before it shows up in referral traffic. And the stakes of being in those answers keep growing: Gartner projected a 25% decline in traditional search volume by 2026 as this behavior mainstreams.5

The intervention map: one lever per mechanism

MechanismTimescaleYour lever
Parametric memoryMonths (model releases)Consistent third-party descriptions in training-weighted sources; allow GPTBot if you want in
RetrievalDays–weeksAllow OAI-SearchBot; answer-first, extractable pages with stats and citations4
ConsensusWeeks–monthsDescription-consistency audit; earn corroborating mentions that repeat your canonical claims
Phrasing & samplingContinuousPrompt suite built on phrasing families; repeated runs; track mention rate with CIs, not screenshots

None of these levers is exotic. What's new is the discipline of working all four at once — and measuring the output the way the system actually behaves: as a probability, not a position.

Frequently asked questions

No. There is no paid placement inside ChatGPT's organic recommendations today. The shortlist is a function of training data, retrieval and consensus — which is precisely why the levers in this post work, and why early movers in a category compound an advantage.

Usually one of two causes: stale parametric memory (the model learned an old price or feature set and hasn't been retrained past it) or a noisy consensus (third-party sources disagree, so the model synthesizes a muddle). Fix the freshest retrievable sources first — searched answers will correct fastest — then work on the corroborating sources for the long loop.

Think in terms of the decision you're making. Ten runs per prompt distinguishes "rarely mentioned" from "usually mentioned"; detecting a 10-point improvement reliably takes dozens of runs across a suite of prompts. The key is consistency: same prompts, same engines, same cadence, confidence intervals on everything.

Sources & further reading

  1. Brown, T. et al. — "Language Models are Few-Shot Learners", arXiv:2005.14165 — documents the GPT-3 training mix (Common Crawl, WebText2, books, Wikipedia) and its weighting.
  2. OpenAI — "Overview of OpenAI crawlers" — GPTBot, OAI-SearchBot and ChatGPT-User, and how to control each.
  3. OpenAI — "OpenAI and Reddit Partnership", May 2024.
  4. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A. — "GEO: Generative Engine Optimization", KDD 2024 / arXiv:2311.09735.
  5. Gartner — "Gartner Predicts Search Engine Volume Will Drop 25% by 2026", February 2024.
Share
ML
Maya Lindqvist

Head of Research at MentionBeat. Maya leads the measurement methodology behind MentionBeat's visibility metrics — prompt-suite design, sampling, and confidence intervals — and writes about how generative engines choose what to say.

Know where you stand in AI answers

MentionBeat samples real buyer prompts across ChatGPT, Claude, Gemini and Perplexity — and turns them into metrics you can act on.

Get your free visibility report
No credit card. Results in about a minute.