LiftInAI. / AI search visibility

Method · published in full, including the prompt set

How I measure AI search visibility

AI search visibility is measured by running a frozen set of buyer-intent prompts repeatedly across several engines and reporting mentions and citations as confidence intervals.

This page exists so you can check my work, argue with it, or copy it outright. None of it is proprietary — the value is in doing it consistently, not in keeping it secret.

If you are comparing consultants, this is the page to ask them for. Any measurement claim that cannot state its prompt set, run count, engine list and reporting rule is not a measurement — it is a screenshot with a number on it.

Book a $1,500 audit Jump to the reference prompt set ↓

The protocol in one card

Prompts
50–200, frozen, six buying stages
Runs
3+ per prompt, same session window
Engines
ChatGPT, Claude, Perplexity, AI Overviews
Counted
Mentions and citations, separately
Reported
Per engine, 95% confidence intervals
Decision rule
Intervals overlap → no trend
Not claimed
Revenue attribution

The measurement problem

Why is measuring this hard in the first place?

Because the same question does not return the same answer. In an independent study of 2,961 repeat runs, the chance of getting an identical list of recommended vendors twice was below 1 in 100, and the chance of an identical order below 1 in 1,000. The variation is not measurement error you can average away by being careful — it is a property of how these systems generate text.

This has a blunt consequence for tooling. Any product that shows you a single “position in AI search” is reporting one draw from a distribution as if it were a fact. That number will move next week whether or not anyone does any work, and it cannot tell you which of those two things happened. Every design choice below exists to get around that one problem.

Repeat runs analysed in the study 2,961
Chance of an identical vendor list twice under 1%
Chance of an identical order twice under 0.1%
Monthly rotation in the domains engines cite 40–60%

Independent measurements. Sources, sample sizes and dates in /research/.

How is the prompt set built?

Pick a buying stage

Not by keyword volume — buyers do not type keywords at an assistant, they describe situations. The set is sampled across six buying stages, because visibility at the top of the funnel and visibility at the vendor-selection stage are different problems with different fixes. A set weighted entirely toward “best tools for X” looks impressive and tells you almost nothing about why you lose in the comparison stage.

What the buyer is doing

The buyer asks the assistant to name options. This is the shortlist moment, and the single most valuable segment of the set.

Example prompts

“best project management tools for creative agencies” · “top alternatives to the market leader for teams under 50”

Share of the set

30%

Weighted highest, because being absent here is the failure mode that costs deals.

How many prompts, and how many runs each?

This is the one design decision that determines whether your report means anything, and it is a genuine trade-off rather than a matter of thoroughness. Every answer collected costs time and money, so assume a fixed budget of 600 answers and look at what different splits buy you. Width beats depth for share of voice; depth is only worth it when you need to diagnose one specific question.

Split of a 600-answer budget Margin of error Good for What it hides
20 prompts × 30 runs ±15.4 pp Diagnosing exactly how one high-value question varies Everything outside those 20 questions — the interval is far too wide to compare quarters
60 prompts × 10 runs ±9.1 pp A mid-size category with a single buyer role Stage-level detail; the interval still swallows most realistic quarterly movement
120 prompts × 5 runs ±6.8 pp A reasonable compromise when a category is narrow Little, but you lose per-stage precision at the edges of the funnel
200 prompts × 3 runs ±5.4 pp Share of voice you can compare against competitors and against next quarter Per-question variance — you see the distribution's shape, not one answer's behaviour

This is why an audit runs 150–200 prompts three times rather than 20 prompts thirty times. The wide split gives a share-of-voice figure precise enough to compare against a competitor and against next quarter; the deep split is reserved for diagnosing a single high-value question where you need to know exactly how the answer varies.

Which engines, and why never a blended number?

Four as standard: ChatGPT, Claude, Perplexity and Google AI Overviews. They are reported as four separate numbers, and I will not average them into one score, because they disagree far more than most dashboards imply — only about 11% of cited domains appear in both ChatGPT and Perplexity. A blended figure would be an average of four different worlds, and it would move for reasons you could never trace.

The practical consequence is that fixes are engine-specific. A citation problem in Perplexity is usually a third-party-source problem; the same symptom in AI Overviews is more often an access or rendering problem. If you only ever see one combined score, you cannot tell which of those you are paying to fix.

Mention or citation — which are you buying?

Both, on separate lines, because they behave differently. In measured samples ChatGPT mentioned a given brand in about 20.7% of relevant answers and attached a citation in roughly 87% of those cases; Gemini's figures were near-mirror at 21.4% and 83.7%. A mention shapes the shortlist. A citation sends a person. Work that raises one does not automatically raise the other.

Are results reproducible for you?

Reproducible in distribution, not per answer. Run my published set yourself and you should land inside the same intervals; you will not land on the same list. That is the honest standard for this medium, and it is the standard I report against.

How does the report read?

Every figure arrives with an interval, and every comparison between two cycles is a paired comparison on the same frozen set — never last month's set against this month's. One rule is written into the engagement before it starts: if the intervals overlap, there is no trend, and the report says there is no trend even when that is the month you would rather show progress.

Reported as

21.4% ±5.4

Share of voice with a 95% interval, per engine. The interval is the deliverable, not decoration.

Compared as

Paired

Cycle against cycle on the identical frozen set, so a change in the questions can never masquerade as a change in visibility.

Decided by

No trend

If this cycle's interval overlaps the last one, the report says there is no trend — the most common honest finding in month two.

What is in every cycle report?

Share of voice and citation rate per engine with intervals, a paired comparison against the baseline, the list of domains cited instead of you, any crawler access changes detected, and a one-page recommendation: continue, change, or stop. It is short on purpose — five pages you will read beat forty you will not.

What can't I measure?

Revenue attribution. There is no reliable way to trace pipeline back to an AI answer today: ChatGPT stopped passing a referrer in roughly 72% of sessions, and a third to two thirds of AI traffic lands in GA4 as Direct. So engagements are contracted on process and leading indicators, and anyone showing you clean AI-sourced revenue is showing you a model, not a measurement.

The reference prompt set, published

Download the full CSV →

Below is a genuine extract from the 187-prompt reference set used in a project-management audit, with brand names generalised — six to eight verbatim prompts per stage, out of the counts shown. The complete set is in the CSV, phrasing and all. Take it, adapt it, run it yourself: if you would rather check my numbers than trust them, this is the file that lets you.

1 · Problem aware

28 prompts in the set

“our team keeps missing client deadlines, what should we change” · “how do agencies keep track of work across many clients” · “we manage projects in spreadsheets and it is falling apart, what now” · “how do I stop status update meetings eating a whole day a week” · “my designers and account managers never see the same priorities” · “what causes scope creep on retainer clients and how is it controlled” · “is there software for knowing whether a project is profitable before it ends” · “how do small studios handle client approvals without endless email”

2 · Category research

28 prompts in the set

“what is project management software and what does it typically cost” · “what should a small agency look for in a project management tool” · “difference between task management and project management software” · “do agencies need resource management or just task tracking” · “what does a per-seat price actually include in PM tools” · “which PM features matter for billable-hours businesses” · “is time tracking usually built in or a separate purchase” · “what integrations does an agency stack normally need”

3 · Vendor discovery

56 prompts in the set

“best project management tools for creative agencies” · “top alternatives to the market leader for teams under 50” · “best PM software for client reporting and profitability” · “which project management tools handle retainers well” · “recommend project management software for a 30-person agency” · “best project tools with built-in time tracking for agencies” · “what do design studios use to manage client work” · “which PM platforms are worth it for under $20 per user”

4 · Comparison

37 prompts in the set

“vendor A vs vendor B for client reporting” · “which of these handles time tracking better, vendor A or vendor C” · “vendor A vs vendor B pricing for 40 seats” · “compare vendor A and vendor D for resource planning” · “is vendor B better than vendor A for agencies specifically” · “vendor A vs spreadsheets, when is it worth switching” · “which is easier to onboard, vendor A or vendor C” · “vendor A vs vendor B for client-facing dashboards”

5 · Validation

22 prompts in the set

“problems with vendor A for agencies” · “is vendor A hard to migrate away from” · “vendor A complaints about billing and seat counts” · “what do users dislike about vendor A after a year” · “does vendor A slow down with large client accounts” · “vendor A support quality for small teams”

6 · Procurement

16 prompts in the set

“does vendor A offer SSO and SOC 2 on the standard plan” · “vendor A annual contract terms for 40 seats” · “is vendor A GDPR compliant and where is data hosted” · “can vendor A be invoiced annually in EUR” · “does vendor A have a DPA and sub-processor list”

Questions about the method

Challenge something here →

If you find a hole in any of this, tell me — I would rather correct the protocol than defend it.

Because one run is a single draw from a distribution and cannot be distinguished from luck. Three runs per prompt is the minimum at which a share-of-voice figure across a wide set becomes stable enough to report with an interval.

I use tools for collection where they help, but not for the headline number. Most report a point estimate from a small set with an undisclosed run count, which is precisely the practice this protocol exists to avoid.

Yes, and the reference prompt set on this page is published so you can. The hard parts are keeping the set frozen, coding mentions and citations consistently, and resisting the urge to report the flattering run.

A mention is the brand named in the answer text; a citation is a link or source attribution pointing at a domain. They are coded separately on every answer, because raising one does not raise the other.

A move whose interval does not overlap the baseline interval on the same engine and the same frozen set. Anything inside the overlap is reported as no change, regardless of how the point estimate looks.

Only in the validation and procurement segments, and never as a headline metric. Branded prompts flatter everyone: engines will describe a company that asks about itself, which tells you nothing about the shortlist.

Because the referrer is largely gone — ChatGPT stopped passing one in roughly 72% of sessions and a third to two thirds of AI traffic appears as Direct in GA4. I contract on process and leading indicators and say so in writing rather than modelling a number and calling it attribution.

Run this on your own category.

Either hand this page to your team and do it yourself, or send me your domain and three competitors and I will run it properly in ten working days. Both outcomes are fine by me; only one of them is billable.

Book a $1,500 audit See what everything costs →
Protocol
Version 3, September 2026
Prompt set
CSV, 187 prompts
Raw data
Published per study
Licence
Copy it freely, credit appreciated
Email
hello@liftinai.com