Case studies of successful LLM-visibility improvement
A curated, credibility-rated evidence base of organizations that measurably improved how often they are cited, mentioned, and recommended in ChatGPT, Perplexity, Gemini, Claude & Google AI Overviews — filtered for what actually transfers to a niche B2B instrument maker. Companion to the LLM Visibility Master Playbook.
GEO is a young field. The overwhelming majority of published case studies are self-reported by the agency or tool vendor that sold the service, with no independent audit, no control group, and survivorship bias (nobody publishes the campaign that flopped). They are still useful — but as evidence of what teams tried, not proof of effect size. To keep this honest, every case below carries a credibility tier.
Credibility tiers used in this document
Outcome metrics in focus (per your brief): citation / mention rate, AI referral traffic & leads, and share of voice vs. competitors.
1The one controlled study — start every argument here
Everything else is anecdote; this is the closest thing to evidence. If you cite one source internally, cite this one.
“GEO: Generative Engine Optimization” (Aggarwal et al., KDD ’24)
A controlled study (randomized + quasi-experimental designs) over GEO-Bench — ~10,000 diverse queries across domains, measuring which on-page content changes increase the probability of being cited in an AI-generated answer. It is the paper that coined the term “GEO.”
What moved the needle (and what didn't)
| Tactic tested | Effect on AI visibility | Read-across |
|---|---|---|
| Add cited statistics (concrete numbers) | strongest (~+41%) | Specs, field-hours, standards numbers in every section |
| Cite authoritative sources | ~+30%; up to ~+115% for lower-ranked pages | The underdog lever — biggest for niche/low-authority topics |
| Add direct quotations from authorities | ~+28% | Quote IEC / WMO / peer-reviewed sources verbatim |
| Improve fluency / clarity | ~+15–28% | Plain, answer-first prose |
| Keyword stuffing | no gain (can hurt) | Don't. Write for intent. |
Across methods, the paper also reports organic-style visibility gains in the ~12–26% range, with efficacy varying by domain — i.e. there is no universal recipe; technical domains need technical evidence.
Sources: arXiv 2311.09735 · ACM SIGKDD KDD’24 · Search Engine Land summary
2Closest analogs — B2B, industrial & technical
The most transferable cases: technical buyers, long sales cycles, authority-driven verticals. All are vendor-reported, so weigh the method over the headline number.
Chemours — authority-led citation in a highly technical vertical
What they did: deep E-E-A-T — content authored by named, recognized industry experts; technical papers that cite patents and peer-reviewed research; and an authoritative backlink profile from industry-specific publications. This is the closest published analog to a scientific-instrument manufacturer.
The most fully documented before / after
| Metric | Before | After | Change | Window |
|---|---|---|---|---|
| AI referral sessions | 43 | 86 | ~+100% | Apr–Sep 2025 |
| AI Overview keywords | 75 | 311 | ~+315% | same |
| Page-1 organic keywords | 29 | 75 | ~+159% | 3–4 mo |
| Qualified leads | — | — | +300% | multi-month |
What they did (maps almost 1:1 to your checklist): answer-first TL;DR blocks, FAQ schema, llms.txt, Q&A-structured content, named expert author bios, consistent monthly publishing, plus earned authority (TechCrunch / G2 / Crunchbase) and a Wikipedia page.
Baselines are tiny — “+100%” is 43→86 sessions. Dramatic percentages, small absolute volume. This is typical of early-stage GEO results and exactly why §6 matters.
“CITABLE” 90-day program
Useful as a 90-day program template. Numbers are agency-reported.
Source: Discovered Labs — B2B SaaS, 3× citations in 90 days.
IEEE Spectrum
Deep, long-form technical content drove a surge in ChatGPT referrals (94% of its AI mix), with referrals still climbing month-over-month in 2025. Closest content-strategy analog to a technical authority brand.
Source: Digiday · RebelMouse.
3Cross-sector wins with transferable lessons
Different industries, but the levers transfer. The recurring pattern — definitional/answer-first content, entity clarity, structured data, third-party authority — is the same spine as your playbook.
| Brand / sector | Reported result | Primary lever | Tier |
|---|---|---|---|
| HubSpot SaaS | Cited in AI Overviews for 3,000+ marketing queries | Years of definitional, entity-first “What is X” content | T2 via Semrush |
| Go Fish Digital Agency (self-test) | +43% AI referral traffic · +83% AI conversions | Prompt-mapping → 5–8 cornerstone assets built for fact-density + external authority | T3 Go Fish Digital |
| Auto-insurance brand Insurance | +447% AI Overview mentions (6 mo) | Structured content, entity clarity, quotable insights | T3 DigitalAgencyNetwork |
| LS Building Products Building materials | +540% AI Overview mentions · +67% organic | Rebuilt content architecture around AI-friendly structure | T3 DigitalAgencyNetwork |
| Farringdons Creative / web | +140% AI traffic · +62% AI mentions | LLM-optimized content + entity reinforcement | T3 DigitalAgencyNetwork |
Strip out the vendor numbers and the same four levers remain in every winning case: (1) answer-first / definitional content, (2) entity clarity, (3) structured data (FAQ / schema), (4) third-party authority. That convergence — across rigorous and vendor sources alike — is the part worth trusting.
4The business case — why visibility is worth the work
Aggregate, mostly third-party data on AI referral traffic and conversion quality. Use these to justify the program; treat precise figures as directional.
Conversion quality T1/T2
AI-referred visitors convert far above organic. A peer-reviewed Marketing Science study plus analytics vendors put ChatGPT at ~16%, Perplexity ~10–11%, vs Google organic ~1.8%; AI-chatbot arrivals were ~38% more likely to purchase in retail.
Sources: Marketing Science (INFORMS) · ALM Corp · Digiday.
Traffic growth T2
Outbound ChatGPT referrals to the web grew ~206% in 2025; ChatGPT referrals up ~52% YoY (Sep–Nov 2025). AI traffic overall rose ~7× from early-2024 to mid-2025.
Sources: Digiday · TechCrunch · Superlines.
Where the traffic is T2
In B2B referrals (early 2026): ChatGPT ~63%, Claude ~18%, Gemini ~11%, Perplexity ~7%. ChatGPT still dominates volume, but optimizing for it alone now covers a third less of the landscape than a year ago — multi-engine matters.
Sources: Goodie (117K+ B2B leads) · Goodie 2026 report.
Share of voice is platform-specific — measure per engine
The same brand and query set produces wildly different share-of-voice by engine (one documented set: Perplexity 28–38%, Gemini 12–20%, ChatGPT 10–16%, Claude 3–7%). This is because engines cite different sources — Perplexity leans heavily on Reddit; ChatGPT favors earned media and Wikipedia.
5Findings & learnings — what the evidence establishes
The synthesis. Each finding is a claim the research supports; each row shows the evidence behind it (with its tier), a confidence rating, and the learning — what to actually do. Confidence reflects the strength of the underlying evidence, not how often the claim is repeated.
| Finding | Evidence | Conf. | Learning — what to do |
|---|---|---|---|
| F1 · Structure and evidence-density cause citations — keywords don't. | Princeton controlled study: +41% from statistics, ~+30% from citing sources, +28% from quotations; keyword stuffing gave no gain. T1 | High | Put concrete numbers, specs and standards in every section; quote authorities verbatim; never optimise for keywords. |
| F2 · The lower your current authority, the more GEO helps. | Same study: citing sources lifted lower-ranked pages by up to ~115% — far above already-authoritative pages. T1 | High | A niche instrument has the most to gain. Seed canonical, corroborated facts first — whoever does owns the answer. Prioritise flagship products. |
| F3 · In technical B2B, authority/E-E-A-T is the winning lever. | Chemours: 82–84% citation via named experts + papers citing patents/peer review T3; IEEE Spectrum surge on deep technical content T2; consistent with the study's corroboration finding T1. | Medium | Lead with primary technical evidence and named expert authorship, not marketing copy. Your IEC/WMO/paper assets are the raw material. |
| F4 · One repeatable on-page recipe recurs in every win. | B2B financing (answer-first/TL;DR, FAQ schema, llms.txt, expert bios) T3; HubSpot definitional content → 3,000+ AI Overview queries T2; multiple vendor cases T3; all consistent with F1 T1. | Medium | Apply the playbook must-haves: definition-first opening, key-facts block, FAQ schema, HTML spec tables, entity/sameAs wiring. |
| F5 · AI-referred visitors convert far above organic. | Peer-reviewed Marketing Science study + analytics: ChatGPT ~16%, Perplexity ~10–11% vs Google organic ~1.8%; AI arrivals ~38% likelier to buy. T1 T2 | High | Even small visibility gains are high-value. Use this to justify the investment — quality, not just volume. |
| F6 · ChatGPT leads volume, but the landscape is fragmenting. | B2B referrals early 2026: ChatGPT ~63%, Claude ~18%, Gemini ~11%, Perplexity ~7%; ChatGPT-only now covers a third less than a year ago. T2 | Medium | Optimise multi-engine. Claude's B2B share is now too large to ignore; don't single-platform. |
| F7 · Share of voice is platform-specific. | Same brand + query set varies 5–10× by engine (e.g. Perplexity 28–38% vs Claude 3–7%), because engines cite different source types. T2 T3 | Medium | Measure and report SOV per engine, never as one blended number. Tailor corroboration to where each engine looks. |
| F8 · Most published numbers are inflated by tiny baselines & survivorship. | Methodological: e.g. “+100%” = 43→86 sessions; only the controlled study isolates cause from effect. T1 | High | Trust the direction, not the decimal. Demand absolute numbers before quoting any case externally. |
| F9 · AI influence is partly invisible to analytics. | AI often shapes a purchase without sending a click; referral counts understate, pipeline-attribution claims overstate. T2 | Medium | Don't judge success on referral clicks alone. Track citation/mention rate & spec-accuracy directly via the prompt suite. |
The four universal levers (the “so-what” of F1–F4)
- Be extractable — answer-first, definition-led, structured HTML, FAQ schema
- Be dense with evidence — concrete statistics, specs, standards in every section (the #1 lever in the controlled study)
- Be corroborated — independent papers, standards bodies, named expert authors, Wikipedia/Wikidata
- Be measured per engine — versioned prompt suite, SOV tracked per platform, re-run on a cadence
The single most important learning (F2)
The controlled study's standout result — citing sources lifts lower-ranked pages by up to ~115%, far more than authoritative ones — flips the usual disadvantage. For a niche instrument with thin search volume and sparse model knowledge, whoever establishes the canonical, corroborated facts first effectively owns the answer. A durable, winnable position, not a race against a famous incumbent.
Confident (Tier-1 backed): structure + cited statistics + quotations increase AI citation; the effect is largest for low-authority pages; AI traffic converts far better than organic. Directional only: the specific uplift percentages from agency cases, exact per-engine share figures, and dollar-pipeline claims — these tell you what to try, not how much you'll get.
6How to read these numbers — the honest caveats
Apply this filter to every case study you encounter (including the ones above).
- Selection bias. Nobody publishes the campaign that flopped. Every percentage is a best case.
- Tiny baselines. “+315%” often means “75→311 keywords.” Always ask for absolute numbers.
- Confounded with ordinary SEO. Most “AI Overview” wins are partly just good SEO, since AI Overviews lean on the existing search index.
- Attribution is hard. AI assistants often influence a purchase without sending a click (“dark” influence) — so referral counts understate impact while pipeline-attribution claims overstate precision.
- Only the Princeton study isolates cause and effect. Use it as the backbone; use the rest as illustrations of what teams tried.
- Figures churn quarterly. Per your playbook's own note: optimize for the direction (structure helps, corroboration helps, entities help), not the decimal.
7What this means for Vaisala
The evidence is unusually well-aligned with the strategy already in your playbook. Three takeaways to carry into the WindCube work and the portfolio rollout.
1 · You're the underdog — that's good
The controlled evidence says niche, lower-authority pages gain most from citing sources and adding statistics. Vaisala's deep technical evidence base (IEC, WMO, papers, field data) is exactly the raw material that wins here.
2 · Evidence > copy
The closest analog (Chemours, technical B2B) won on named expertise + primary research + standards, not marketing language. Lead with specs, standards, and cited numbers — the SSOT discipline you already enforce.
3 · Measure per engine, prove direction
SOV differs 5–10× between Perplexity and Claude for the same query. Stand up the versioned prompt suite (§8 of the playbook), baseline now, and report per-engine citation rate, SOV, and spec-accuracy — not a blended figure.
No public case study is a perfect proxy for a scientific-instrument maker, and the rigorous evidence is thin — but it all points one way, and that way is the strategy you've already written down: be the clearest, most evidence-dense, most corroborated true source about your product, and measure it per engine. The case studies don't change the plan; they justify the investment.
§Sources & references
Grouped by credibility tier. Vendor-reported figures (Tier 3) should be cited internally with that caveat attached.
Tier 1 — rigorous / peer-reviewed
- Aggarwal et al., “GEO: Generative Engine Optimization” — arXiv 2311.09735 · ACM SIGKDD KDD’24
- Search Engine Land, GEO framework introduced in research paper
- Marketing Science (INFORMS), ChatGPT referrals to e-commerce vs. traditional channels
Tier 2 — independent analytics / analyst data
- Digiday, State of AI referral traffic in 2025 · ChatGPT = 20% of Walmart referral traffic
- ALM Corp, ChatGPT converts 31% higher than non-branded organic
- Goodie, AI search vs. traditional search — 117K+ B2B leads · 2026 AI Search Traffic Report
- Superlines, AI Search Statistics 2026 · TechCrunch, ChatGPT retailer-app referrals +28% YoY
- Semrush, GEO practical guide (HubSpot AI Overviews analysis)
Tier 3 — vendor / agency self-reported (method > numbers)
- DigitalAgencyNetwork, GEO case studies roundup (Chemours, auto-insurance, LS Building, Farringdons, HubSpot)
- Concurate, B2B financing platform GEO case study
- Discovered Labs, B2B SaaS 3× citation rate in 90 days
- Go Fish Digital, GEO case study: driving leads
- RebelMouse, IEEE Spectrum ChatGPT referral growth
Compiled June 2026 via live web research. Figures reflect what each source reported at time of access; AI-visibility metrics shift quarterly — re-verify before quoting externally. Optimize for the direction, not the decimal.