- Starting point: 9% mention rate on core buying prompts, frequently misdescribed, absent from cited sources.
- Three interventions: answer-first page rewrites, comparison pages, and a corroboration campaign.
- Result: 38% mention rate and a jump in recommendation rate on ‘which should I buy’ prompts — all measured against a baseline with confidence intervals.
- The lever wasn’t volume; it was accuracy, structure and corroboration.1
The brand in this example — call it a mid-market project-management SaaS — had strong SEO and a decent product, but was nearly invisible in AI answers. When buyers asked assistants for recommendations in its category, it appeared about one time in ten, often with a stale feature description, while two rivals split the rest. This is a composite drawn from recurring patterns, not a single named client, but every number reflects the kind of movement these interventions produce.
Week 1 — The baseline
The team built a suite of 40 buyer prompts spanning discovery (“best tools for X”), comparison (“X vs Y”), and recommendation (“which should a 20-person team pick”). They ran each across four engines, multiple times, and computed rates with confidence intervals. The baseline was sobering:
- Mention rate: 9% (±3) on core buying prompts.
- Recommendation rate: 4% — named occasionally, recommended almost never.
- Accuracy: two recurring errors — a discontinued limitation and a wrong integration claim.
- Cited sources: reviews and a key subreddit cited rivals; the brand was absent.
The baseline did more than set a number — it diagnosed the problem. Low mention rate plus recurring errors plus absence from cited sources pointed at three specific fixes, not a vague ‘do more content.’
Weeks 2–5 — The interventions
1. Answer-first page rewrites
The team rewrote its highest-stakes pages to lead with a direct, self-contained answer to each page’s buyer question, replaced superlatives with specific sourced claims, and added Product and FAQ schema. This targeted both quotability and the two accuracy errors.
2. Comparison pages
They published fair, structured “vs” pages against the two rivals that were winning — leading with a verdict, a factual table (including where the rivals were better), and “choose us if…” guidance. These mapped directly onto the comparison prompts in the suite.
3. A corroboration campaign
Finally, they earned genuine third-party coverage and reviews on the exact sites the engines had been citing — the review platform and the community thread where rivals appeared and they didn’t — so consensus began to include them.
| Metric | Baseline | Week 12 |
|---|---|---|
| Mention rate (buying prompts) | 9% (±3) | 38% (±5) |
| Recommendation rate | 4% | 19% |
| Recurring factual errors | 2 | 0 |
| Cited-source presence | Absent | Present on 2 key domains |
Week 12 — The re-measurement
Re-running the identical suite twelve weeks later, mention rate on core buying prompts had risen from 9% to 38% — a change well beyond the confidence bands, so not noise. Recommendation rate nearly quintupled off a low base. Both accuracy errors were gone, and the brand now appeared in the cited sources that had previously featured only rivals. The gains concentrated exactly where the interventions were aimed: comparison and recommendation prompts.
The brand didn’t publish more than its rivals. It published the specific, accurate, corroborated things the model needed to include it.
What generalizes
- Baseline first — the diagnosis is in the breakdown, not the headline number.
- Fix accuracy early — wrong claims disqualify you before quotability even matters.
- Build comparison pages against whoever is actually winning your prompts.
- Earn corroboration on the exact sources the engines already cite.
- Re-measure against baseline with intervals — attribute lift to the work, not to luck.
Frequently asked questions
It’s a composite, illustrative example built from patterns we see repeatedly, not a single named customer — the numbers reflect the kind of movement these interventions typically produce. It’s framed that way deliberately, because inventing a specific client with precise figures would be dishonest.
Three compounding fixes: pages the model could quote accurately, comparison content matching the exact buying prompts, and corroboration on the sources the engines already cited. No single tactic did it — the combination is what moved the number beyond noise.
The retrieval-driven gains (comparison pages, corrected facts) can show within weeks of indexing; corroboration compounds over a quarter. Most programs see meaningful movement on targeted prompts inside 8–12 weeks.
Sources & further reading
- "GEO: Generative Engine Optimization", Aggarwal et al., KDD 2024 / arXiv:2311.09735.
- Pew Research Center — "Google users are less likely to click on links when an AI summary appears", July 2025.
- Gartner — "Search Engine Volume Will Drop 25% by 2026", February 2024.
- Schema.org vocabulary — Product, Offer, FAQPage, Organization types.