MentionBeat · Knowledge base

How MentionBeat measures AI visibility

Exactly how our numbers are produced — including what we do not do. Every claim on this page corresponds to shipped, testable code. Most tools in this category publish no methodology; we think the method is the product.

Last updated 2026-07-10 · Playbook · Case studies

1Collection: real engines, real answers

Every measurement response comes from the real engine's own API, called with your tracked prompt at a realistic sampling temperature. We do not infer AI answers from Google rankings, and we never silently substitute one model for another.

EngineHow we query itWeb-grounded?
ChatGPTOpenAI APIOptional (OpenAI web search)
ClaudeAnthropic APIOptional (Anthropic web-search tool)
GeminiGoogle Gemini APIOptional (Google Search grounding)
PerplexityPerplexity Sonar APIAlways (natively grounded)
GrokxAI APIOptional (Live Search over web + X)
DeepSeekDeepSeek APINo — parametric only, and we refuse to label it otherwise
Google AI OverviewsSERP capture (SerpApi)Always
Google AI ModeSERP capture (SerpApi)Always
Microsoft Copilot · Meta AINot yet measured. No API or SERP surface exists; a browser-capture adapter is on the roadmap. We'd rather show a gap than simulated data.

2Sampling: repetition and honesty about variance

LLMs are non-deterministic; a single query is an anecdote, not a measurement.

Anti-leakage design
Headline visibility uses only unbranded prompts — questions a real buyer would ask before knowing your name. Brand-bearing prompts are tracked but reported separately and excluded from headline share-of-voice. Asking an engine “tell me about Acme” and counting the reply as visibility is grading your own exam.

3Scoring: a calibrated judge, not keyword matching

Responses are scored by an LLM judge that extracts brand mentions, position, sentiment, recommendation strength, cited domains and factual claims. Because judges are also LLMs, we treat the judge itself as an instrument to calibrate:

4Statistics: confidence intervals or it didn't happen

5Accuracy intelligence: grounded in your approved facts

We don't just count mentions — we fact-check what engines say about you against your single source of truth (SSOT): a provenance-tracked store of approved product facts, each with its source and human-review status.

6Experiments: measured lift, not vibes

Content changes are evaluated as baseline → treatment experiments: per-metric lift with confidence intervals, significant only when intervals separate. A confound guard flags experiments where the model set, repetitions, profile version, or provider model version changed mid-experiment — those are marked confounded instead of quietly reported as wins.

7Live / modeled honesty

Labeled, end to end
The authoring studio can preview engine behavior through a persona simulation when no API key is configured. Every such run is labeled modeled (or mixed). Modeled data never appears in measurement reporting as if it were live — if a vendor key is missing you see a gap or a label, never an imitation.

8Known limitations (the part most vendors skip)