Methodology & limitations

Judge the questions before the answers

Versioned prompt library, clean logged-out sessions, per-engine scoring, three runs on separate days for audits, ambiguous cases scored against my own client, raw data shipped with every report. The full methodology - including its limitations - is public. Judge the questions before the answers.

Three core engines, one observed surface

Every engagement measures ChatGPT, Gemini and Perplexity - the three systems B2B buyers most commonly ask during vendor research.

Google AI Overviews appears only in the Audit, as an observed section outside the scorecard: its answers are recorded and annotated, but not scored, because its retrieval behavior differs too much from conversational engines to share a scale.

Prompt library, sessions and runs

  • Versioned prompt library. Prompts are written once per category, versioned, and reused unchanged across runs so results stay comparable.
  • Clean logged-out sessions. Every run uses fresh, logged-out sessions - a controlled baseline, deliberately free of personal history.
  • Snapshot runs. One base run across all 15 prompts plus one additional run for the 5 headline prompts (single-run - directional).
  • Audit runs. Three runs on separate days, so day-to-day output variability is measured rather than assumed.

What gets scored, and how strictly

Each engine's answers are evaluated on three observable events:

  • Mention - the brand is named at all.
  • Recommendation - the brand is suggested as an option.
  • Citation - the brand's domain is linked as a source.

The scoring model combines its components with fixed weights:

The mapping of each weight to a named component is published after owner verification; until then only the weights themselves are stated.

  • Ambiguous cases are scored against my own client. If an answer could be read either way, it counts as a miss.
  • Low-confidence outputs are flagged as such in the raw data rather than silently classified.
  • Raw data and annotated screenshots ship with every report - every scored cell traces to a prompt, an engine, a date and the engine's full output.

Four limitations, stated plainly

Outputs vary over time

Engines change their answers between days and model updates. A report is a dated record, not a permanent state. Audits use three runs on separate days to measure this; Snapshots do not.

15 prompts is a sample

A Snapshot covers the highest-signal questions in your category, not every question a buyer could ask. It ranks gaps and priorities; it does not enumerate them exhaustively.

Controlled baseline is not every buyer

Clean logged-out sessions are reproducible, but a real buyer's personalised session may differ. The baseline shows what engines say by default, not what any single person sees.

No business-outcome guarantee

The report diagnoses visibility and misrepresentation and prioritises fixes. It does not and cannot guarantee rankings, traffic, leads or revenue.

Directional but decision-useful: precise enough to rank your gaps and priorities, honest enough to say what a 15-prompt scan can't tell you.

Claims discipline

Third-party statistics appear on this site only after the source is verified and recorded in a claims register with its label, URL and publication date. Unverified claims are not shown at all - there are currently no verified third-party claims published on this page.

Findings reflect engine outputs under the stated protocol on the stated dates; outputs vary over time. This report is diagnostic information, not a guarantee of business outcomes.

Not affiliated with, endorsed by, or partnered with OpenAI, Google, or Perplexity AI.

ChatGPT, Gemini, Perplexity and all product names are trademarks of their respective owners.