- Training corpora are dominated by filtered web crawl plus deliberately upweighted curated sources like Wikipedia — the GPT-3 paper documents this mix explicitly.1
- Licensing deals — OpenAI–Reddit (May 2024)2 and Google–Reddit (Feb 2024)3 — made community discussion first-class model food, not incidental scrapings.
- Retrieval engines lean hard on forums and review sites for "best X" queries, because that's where candid comparative opinion lives.
- The playbook on every platform is the same at its core: be genuinely, verifiably discussed. Astroturfing backfires on all of them — and the platforms' rules and communities actively hunt it.
Marketers persistently overrate how much of a model's opinion comes from their own website. Your domain is one voice in a corpus of billions of documents — and for the questions that decide purchases ("which one is actually good?"), it's a voice the model knows is biased.
So where do the opinions actually come from? Follow the data: what goes into training, what the labs pay to license, and what the retrieval layer fetches when someone asks for a recommendation. Three kinds of places keep showing up — community forums, reference works, and review platforms. This post takes them in turn.
What the training-data record shows
We rarely get full transparency into modern training sets, but the foundational paper is unusually specific. Brown et al.'s GPT-3 paper lists the corpus: a filtered Common Crawl, the WebText2 corpus of link-shared pages, two book corpora, and English Wikipedia. Crucially, the mix was not proportional to size — higher-quality corpora were sampled more often, with Wikipedia and the curated sets seen multiple times during training while the raw crawl was sampled less than once through.1
Two durable lessons from that design choice, which subsequent labs have echoed in spirit if not in published detail:
- Curation beats volume. A paragraph in a heavily-sampled reference source can outweigh dozens of ordinary web pages. If Wikipedia describes your category and you're absent from that description, you're absent from something the model read repeatedly.
- Link-shared and discussed content punches above its weight. WebText-style corpora select pages that real people found worth sharing — which is why forum-endorsed content keeps surfacing in model knowledge.
The licensing deals made it official
For years, community content entered models through scraping — incidental and legally contested. Then the checkbooks came out. In February 2024, Reddit and Google announced a licensing arrangement giving Google access to Reddit's data (via its Data API) for training and products, widely reported at around $60M a year.3 In May 2024, OpenAI announced its own Reddit partnership, bringing Reddit content into OpenAI products via Reddit's API.2
Read those deals as a statement of what the labs believe: authentic human discussion is among the most valuable text on the internet for teaching models what people actually think about products, brands and trade-offs. When a lab pays eight figures a year for r/homelab and r/skincareaddiction, "what Reddit thinks of your brand" stops being a community-management concern and becomes a model-knowledge concern.
Timing note: licensed data flows into products on two clocks. Retrieval features can surface fresh threads within days; training ingestion shows up model-generations later. A thread written today can influence answers this week via search — and next year via weights.
Why "best X" retrieval loves forums
Ask Perplexity or ChatGPT's search mode for "best standing desk under $500" and check the citations: alongside a couple of publisher roundups you'll almost always find Reddit threads and review-site pages. That's not an accident of ranking — it's fit-for-purpose. A "best X" query is a request for comparative experience, and forums are the only place on the web where large numbers of people compare products with no commercial stake, in exactly the vocabulary buyers use.
Sampling shortlist-style queries across engines, we consistently see citation mixes that look roughly like this — the proportions below are illustrative, not a measured benchmark, but the shape will be familiar to anyone who checks their own category:
The strategic reading: for recommendation queries, roughly half the evidence an assistant weighs is community and review content — places you can participate in but not control. Your site supplies specs and pricing; everyone else supplies the verdict.
This also explains a pattern that puzzles teams new to GEO: a brand with immaculate documentation and strong SEO that barely appears in shortlist answers. Their owned content wins the informational queries ("how does X work") but the recommendation queries route around it, because the engine is deliberately looking for what other people say. If your category subreddit hasn't mentioned you in eighteen months and your newest G2 review predates your last two releases, the assistant is synthesizing a verdict from a version of you that no longer exists — or skipping you entirely.
Reddit: participate, don't perform
Reddit rewards exactly one strategy, and it's the slow one.
Genuine participation. Have named employees — flaired, disclosed — answer technical questions in your category subreddits. Not product pitches: actual answers, including "our product isn't the right fit for that" when it isn't. Over months, this produces the threads that retrieval later cites: a real person from the company being competent in public. Founder and engineer accounts consistently do better than brand accounts, because expertise is the currency.
AMAs done properly. A well-run AMA in a relevant subreddit — coordinated with moderators, with a technical founder answering hard questions candidly — creates a dense, permanent, highly-linked document about your brand in conversational Q&A form. That format is nearly ideal training and retrieval material.
Monitoring as measurement. Track threads where your category's buying questions get asked ("what do people use for X these days?"). These threads are leading indicators of what assistants will say next quarter — and when a factual error about your product spreads in one, a polite, disclosed correction from a company account is both good community practice and good GEO.
- Disclose affiliation in every comment; use flair where subreddits offer it.
- Answer questions where you're the genuine expert — including unfavorable ones.
- Work with moderators before AMAs or any promotional activity.
- Correct factual errors about your product politely, with sources.
- Run sockpuppets, buy upvotes, or seed fake "just a happy customer" threads.
- Ask employees to post "organic" recommendations without disclosure.
- Blitz multiple subreddits with the same talking points.
- Delete critical threads' rebuttals or argue with reviewers.
To be blunt about the "don't" column: astroturfing backfires, specifically and durably. It violates Reddit's rules, communities are unnervingly good at unmasking it, and the unmasking thread — "Company X caught posting fake reviews" — is itself high-engagement content that gets linked, upvoted, and eventually learned by the same models you were trying to game. You'd be training the corpus on your own scandal.
MentionBeat runs real "best X" and comparison prompts across ChatGPT, Claude, Gemini and Perplexity, shows which sources get cited in your category, and tracks how your mention rate moves as the community record changes.
Get a free visibility reportWikipedia: influence the sources, not the article
Wikipedia's outsized sampling weight in training corpora1 makes it tempting — and it is precisely the platform where marketing instincts do the most damage. Editing your own article, or paying someone to, violates Wikipedia's conflict-of-interest norms; undisclosed paid editing violates its terms of use. Editors are experienced at detecting it, and an article's talk page documenting your COI editing is a permanent, model-readable record of the attempt.
The legitimate route runs one layer upstream. Wikipedia is a mirror: it reflects what independent reliable sources have already published. You don't polish the mirror — you change what stands in front of it:
- Earn notability instead of asserting it. Wikipedia summarizes independent reliable sources. No significant third-party coverage, no article — and no amount of drafting fixes that. Analyst reports, substantial trade-press features and mainstream coverage are the prerequisites.
- Improve the sources Wikipedia cites. Category articles ("time-series database," "electronic lab notebook") cite surveys, textbooks, trade publications. Getting your brand accurately covered in those citable sources is how you legitimately enter the reference layer.
- Use talk pages and edit requests for errors. If your article contains factual mistakes, the sanctioned path is a disclosed {{edit request}} on the talk page with reliable sources attached. It's slower than editing directly. It also actually works.
- Watch, don't touch. Put your article and your category's articles on a monitoring cadence. Changes there propagate into both retrieval answers and future training snapshots.
G2, Capterra, TrustRadius: the structured opinion layer
For B2B software, review platforms are where assistants find the structured version of the community verdict: pros/cons lists, feature ratings, segment fit ("best for mid-market"), all in clean crawlable text. When an assistant hedges — "users praise the onboarding but report a steep learning curve for reporting" — it is very often paraphrasing review-platform content.
What matters, in order:
- Existence and completeness. Claim your profiles on G2, Capterra and TrustRadius; fill in categories, features, pricing and screenshots. An empty profile forfeits the co-occurrence with category terms these pages generate.
- Recency. A wall of glowing 2023 reviews reads as a product in decline — to humans and to models summarizing "current sentiment." Steady review velocity beats occasional spikes; build the ask into onboarding milestones and QBRs, offered to happy and unhappy customers alike (platform rules require it, and mixed-but-fresh beats perfect-but-stale).
- Responses to negative reviews. A composed, specific vendor response is crawlable text that often gets summarized alongside the complaint — it's your only sanctioned way to add context to criticism.
- Consistency with your positioning. Your profile description should use the same positioning phrase as your site and press boilerplate, so every surface reinforces one association.
Then verify the loop end-to-end: sample the assistants' actual answers about your category and see whether the review-layer story is the one being told. This cross-engine sampling is tedious by hand — it's the kind of monitoring a platform like MentionBeat runs continuously — but however you do it, do it before and after you invest, so the effect is measured rather than assumed.
Frequently asked questions
They can, especially if the thread ranks well for retrieval and nothing contradicts it. Your options are all upstream: a disclosed, factual response in the thread itself; newer positive discussion that outweighs it (earned, not manufactured); and fixing the underlying product complaint so the future record diverges from the past one. What doesn't work is deletion requests or vote manipulation — both tend to amplify the original complaint.
Only if independent coverage already supports notability — otherwise the article gets deleted and the deletion log lingers. The better question is usually whether your category's articles describe the space accurately and cite sources where you appear. For most mid-size brands, presence in well-cited category sources beats a thin brand article that's one deletion discussion from vanishing.
Review platforms, usually — they're the cheapest to influence legitimately (you're mobilizing customers you already have) and the most directly cited for B2B shortlist queries. Reddit is the highest ceiling but the slowest compounding. Wikipedia is a consequence of the other two more than a starting point. Whatever you pick, baseline your mention rate first so you can attribute the movement.
Sources & further reading
- Brown, T., et al. — "Language Models are Few-Shot Learners", arXiv:2005.14165 — documents GPT-3's training mix and the upweighted sampling of Wikipedia and curated corpora relative to raw web crawl.
- OpenAI — "OpenAI and Reddit Partnership", May 2024.
- Reddit, Inc. — Reddit newsroom; the Google licensing arrangement was announced in February 2024 and widely reported at roughly $60M/year.
- Aggarwal, P., et al. — "GEO: Generative Engine Optimization", KDD 2024 / arXiv:2311.09735 — on which content signals measurably lift visibility inside generative answers.