Method · published in full, including the prompt set
AI search visibility is measured by running a frozen set of buyer-intent prompts repeatedly across several engines and reporting mentions and citations as confidence intervals.
This page exists so you can check my work, argue with it, or copy it outright. None of it is proprietary — the value is in doing it consistently, not in keeping it secret.
If you are comparing consultants, this is the page to ask them for. Any measurement claim that cannot state its prompt set, run count, engine list and reporting rule is not a measurement — it is a screenshot with a number on it.
The protocol in one card
The measurement problem
Because the same question does not return the same answer. In an independent study of 2,961 repeat runs, the chance of getting an identical list of recommended vendors twice was below 1 in 100, and the chance of an identical order below 1 in 1,000. The variation is not measurement error you can average away by being careful — it is a property of how these systems generate text.
This has a blunt consequence for tooling. Any product that shows you a single “position in AI search” is reporting one draw from a distribution as if it were a fact. That number will move next week whether or not anyone does any work, and it cannot tell you which of those two things happened. Every design choice below exists to get around that one problem.
Independent measurements. Sources, sample sizes and dates in /research/.
Pick a buying stage
Not by keyword volume — buyers do not type keywords at an assistant, they describe situations. The set is sampled across six buying stages, because visibility at the top of the funnel and visibility at the vendor-selection stage are different problems with different fixes. A set weighted entirely toward “best tools for X” looks impressive and tells you almost nothing about why you lose in the comparison stage.
What the buyer is doing
The buyer asks the assistant to name options. This is the shortlist moment, and the single most valuable segment of the set.
The buyer has a symptom, not a category. They describe the pain in their own words and let the assistant name the type of tool that solves it.
The buyer now knows the category name and is learning how it works, what it costs, and what distinguishes one option from another.
Two or three named vendors are compared head to head, usually on price, integrations, or a specific capability.
The buyer has a preferred option and is looking for reasons not to buy it: complaints, migration pain, hidden limits.
Commercial and compliance questions from the people who sign rather than the people who use.
Example prompts
“best project management tools for creative agencies” · “top alternatives to the market leader for teams under 50”
“our team keeps missing client deadlines, what should we change” · “how do agencies keep track of work across many clients”
“what is project management software and what does it cost” · “what should a small agency look for in a PM tool”
“vendor A vs vendor B for client reporting” · “which of these handles time tracking better”
“problems with vendor A for agencies” · “is vendor A hard to migrate away from”
“does vendor A offer SSO and SOC 2 on the standard plan” · “vendor A annual contract terms for 40 seats”
Share of the set
30%
15%
15%
20%
12%
8%
Weighted highest, because being absent here is the failure mode that costs deals.
Cheap visibility, low intent. Worth measuring so you know whether you exist at all upstream.
This is where definitional content earns citations rather than mentions.
Where third-party comparison pages dominate the citations almost entirely.
The segment most often ignored, and the one where a bad Reddit thread does real damage.
Small share, high consequence: wrong facts here kill a deal that was already won.
This is the one design decision that determines whether your report means anything, and it is a genuine trade-off rather than a matter of thoroughness. Every answer collected costs time and money, so assume a fixed budget of 600 answers and look at what different splits buy you. Width beats depth for share of voice; depth is only worth it when you need to diagnose one specific question.
| Split of a 600-answer budget | Margin of error | Good for | What it hides |
|---|---|---|---|
| 20 prompts × 30 runs | ±15.4 pp | Diagnosing exactly how one high-value question varies | Everything outside those 20 questions — the interval is far too wide to compare quarters |
| 60 prompts × 10 runs | ±9.1 pp | A mid-size category with a single buyer role | Stage-level detail; the interval still swallows most realistic quarterly movement |
| 120 prompts × 5 runs | ±6.8 pp | A reasonable compromise when a category is narrow | Little, but you lose per-stage precision at the edges of the funnel |
| 200 prompts × 3 runs | ±5.4 pp | Share of voice you can compare against competitors and against next quarter | Per-question variance — you see the distribution's shape, not one answer's behaviour |
This is why an audit runs 150–200 prompts three times rather than 20 prompts thirty times. The wide split gives a share-of-voice figure precise enough to compare against a competitor and against next quarter; the deep split is reserved for diagnosing a single high-value question where you need to know exactly how the answer varies.
Four as standard: ChatGPT, Claude, Perplexity and Google AI Overviews. They are reported as four separate numbers, and I will not average them into one score, because they disagree far more than most dashboards imply — only about 11% of cited domains appear in both ChatGPT and Perplexity. A blended figure would be an average of four different worlds, and it would move for reasons you could never trace.
The practical consequence is that fixes are engine-specific. A citation problem in Perplexity is usually a third-party-source problem; the same symptom in AI Overviews is more often an access or rendering problem. If you only ever see one combined score, you cannot tell which of those you are paying to fix.
Both, on separate lines, because they behave differently. In measured samples ChatGPT mentioned a given brand in about 20.7% of relevant answers and attached a citation in roughly 87% of those cases; Gemini's figures were near-mirror at 21.4% and 83.7%. A mention shapes the shortlist. A citation sends a person. Work that raises one does not automatically raise the other.
Reproducible in distribution, not per answer. Run my published set yourself and you should land inside the same intervals; you will not land on the same list. That is the honest standard for this medium, and it is the standard I report against.
Every figure arrives with an interval, and every comparison between two cycles is a paired comparison on the same frozen set — never last month's set against this month's. One rule is written into the engagement before it starts: if the intervals overlap, there is no trend, and the report says there is no trend even when that is the month you would rather show progress.
Reported as
21.4% ±5.4
Share of voice with a 95% interval, per engine. The interval is the deliverable, not decoration.
Compared as
Paired
Cycle against cycle on the identical frozen set, so a change in the questions can never masquerade as a change in visibility.
Decided by
No trend
If this cycle's interval overlaps the last one, the report says there is no trend — the most common honest finding in month two.
Share of voice and citation rate per engine with intervals, a paired comparison against the baseline, the list of domains cited instead of you, any crawler access changes detected, and a one-page recommendation: continue, change, or stop. It is short on purpose — five pages you will read beat forty you will not.
Revenue attribution. There is no reliable way to trace pipeline back to an AI answer today: ChatGPT stopped passing a referrer in roughly 72% of sessions, and a third to two thirds of AI traffic lands in GA4 as Direct. So engagements are contracted on process and leading indicators, and anyone showing you clean AI-sourced revenue is showing you a model, not a measurement.
Below is a genuine extract from the 187-prompt reference set used in a project-management audit, with brand names generalised — six to eight verbatim prompts per stage, out of the counts shown. The complete set is in the CSV, phrasing and all. Take it, adapt it, run it yourself: if you would rather check my numbers than trust them, this is the file that lets you.
1 · Problem aware
28 prompts in the set
“our team keeps missing client deadlines, what should we change” · “how do agencies keep track of work across many clients” · “we manage projects in spreadsheets and it is falling apart, what now” · “how do I stop status update meetings eating a whole day a week” · “my designers and account managers never see the same priorities” · “what causes scope creep on retainer clients and how is it controlled” · “is there software for knowing whether a project is profitable before it ends” · “how do small studios handle client approvals without endless email”
2 · Category research
28 prompts in the set
“what is project management software and what does it typically cost” · “what should a small agency look for in a project management tool” · “difference between task management and project management software” · “do agencies need resource management or just task tracking” · “what does a per-seat price actually include in PM tools” · “which PM features matter for billable-hours businesses” · “is time tracking usually built in or a separate purchase” · “what integrations does an agency stack normally need”
3 · Vendor discovery
56 prompts in the set
“best project management tools for creative agencies” · “top alternatives to the market leader for teams under 50” · “best PM software for client reporting and profitability” · “which project management tools handle retainers well” · “recommend project management software for a 30-person agency” · “best project tools with built-in time tracking for agencies” · “what do design studios use to manage client work” · “which PM platforms are worth it for under $20 per user”
4 · Comparison
37 prompts in the set
“vendor A vs vendor B for client reporting” · “which of these handles time tracking better, vendor A or vendor C” · “vendor A vs vendor B pricing for 40 seats” · “compare vendor A and vendor D for resource planning” · “is vendor B better than vendor A for agencies specifically” · “vendor A vs spreadsheets, when is it worth switching” · “which is easier to onboard, vendor A or vendor C” · “vendor A vs vendor B for client-facing dashboards”
5 · Validation
22 prompts in the set
“problems with vendor A for agencies” · “is vendor A hard to migrate away from” · “vendor A complaints about billing and seat counts” · “what do users dislike about vendor A after a year” · “does vendor A slow down with large client accounts” · “vendor A support quality for small teams”
6 · Procurement
16 prompts in the set
“does vendor A offer SSO and SOC 2 on the standard plan” · “vendor A annual contract terms for 40 seats” · “is vendor A GDPR compliant and where is data hosted” · “can vendor A be invoiced annually in EUR” · “does vendor A have a DPA and sub-processor list”
If you find a hole in any of this, tell me — I would rather correct the protocol than defend it.
Because one run is a single draw from a distribution and cannot be distinguished from luck. Three runs per prompt is the minimum at which a share-of-voice figure across a wide set becomes stable enough to report with an interval.
I use tools for collection where they help, but not for the headline number. Most report a point estimate from a small set with an undisclosed run count, which is precisely the practice this protocol exists to avoid.
Yes, and the reference prompt set on this page is published so you can. The hard parts are keeping the set frozen, coding mentions and citations consistently, and resisting the urge to report the flattering run.
A mention is the brand named in the answer text; a citation is a link or source attribution pointing at a domain. They are coded separately on every answer, because raising one does not raise the other.
A move whose interval does not overlap the baseline interval on the same engine and the same frozen set. Anything inside the overlap is reported as no change, regardless of how the point estimate looks.
Only in the validation and procurement segments, and never as a headline metric. Branded prompts flatter everyone: engines will describe a company that asks about itself, which tells you nothing about the shortlist.
Because the referrer is largely gone — ChatGPT stopped passing one in roughly 72% of sessions and a third to two thirds of AI traffic appears as Direct in GA4. I contract on process and leading indicators and say so in writing rather than modelling a number and calling it attribution.
Either hand this page to your team and do it yourself, or send me your domain and three competitors and I will run it properly in ten working days. Both outcomes are fine by me; only one of them is billable.