Method · last revised 26 August 2026
How we measure whether an AI assistant recommends a business
This page exists so that every number in one of our reports can be checked, argued with, or reproduced by somebody else. If a figure we give you is not explained here, treat that as our mistake and tell us.
What the measurement actually is
We ask four AI assistants a fixed set of questions a customer might ask, several times each, and record which businesses each answer names. The output is a recommendation rate: the proportion of answers in which a given business appeared. It is not a rank, a score out of a hundred, or a position. Those do not exist inside an AI answer, and a report that offers you one has invented it.
6
prompts per report
3
samples per prompt, per engine
72
answers read in total
Where the prompts come from
Prompts are written in the words a customer would use, not the words a marketer would. Each one has to name a service and a place, describe a situation rather than a product, and be a thing somebody would plausibly type at eleven at night. We draw them from the business’s own service pages, the questions it says it gets asked, and the phrasing already visible in its search terms.
They are agreed with you before the run and frozen for the engagement. Changing the prompts changes the number, so a prompt set that drifts month to month makes every comparison meaningless — which is a convenient way for an agency to show progress that is not there.
Which assistants, which models
Four, with web search enabled and the location set to the town you serve. The exact model identifier is recorded against every answer and printed in the report, because “we asked ChatGPT” is not a reproducible statement.
| Assistant | Model recorded | Web search |
|---|---|---|
| ChatGPT | gpt-5.6 | on |
| Claude | claude-sonnet-5 | on |
| Gemini | gemini-3-pro | on |
| Perplexity | sonar | on |
Model identifiers are those current on the date of the run. When a provider retires one, the change is noted in the report rather than absorbed silently.
Why we ask the same question three times
Because the same question does not return the same answer twice. These systems are non-deterministic: ask “best physiotherapist in Leeds” three times and you can get three different sets of names. A single ask is therefore not a measurement, it is an anecdote, and a report built from one is not reproducible even by the people who produced it.
Three samples is a floor, not a boast. It is enough to distinguish “never named” from “sometimes named”, which is the distinction that matters commercially. It is not enough to put a confidence interval on a small difference, so we do not quote one.
Cheaper is available. Asking each question once, without web search, against the smallest model a provider sells costs roughly a tenth of what we spend. It also measures what a model half-remembers rather than what your customer sees, and a prospect can disprove it in thirty seconds by opening ChatGPT. We do not do that.
What counts as being named
- CountsThe business is named as an option, in any position, in the answer text.
- CountsA recognisable variant of the name — trading name, with or without Ltd, with or without the town.
- Does not countAppearing only in a citation, footnote or source link without being named in the prose. The customer does not read those.
- Does not countBeing named as an example of what to avoid, or in an aside unrelated to the question.
- Does not countA directory, comparison site or marketplace that happens to list the business. That is the directory being recommended, not you.
The map grid, and what the percentage means
We lay 25 points in a five-by-five square over the service area and ask Google, at each one, who it would show a customer standing there. The headline percentage is the share of those points where the business appears in the top three. We also report the mean rank across the points where it appears at all, and the mean across every point with a default applied where it is absent — because the first number alone flatters a business that is either brilliant or missing.
The zoom level of each lookup decides how large an area that point “sees”, and getting it wrong produces a confident, colourful, meaningless picture. Ours is recorded in the report with the coordinates, so the same grid can be re-run.
The on-site checks
These are the only part of the report that is not probabilistic. Each is read directly off your public pages, carries the literal line of evidence rather than a summary of it, and is either true or false — whether a crawler is disallowed in your robots file, whether an llms.txt exists, whether a machine can resolve your identity, address and hours. No password, no access, nothing installed. What our scanner requests, and how to block it.
What this measurement cannot see
What a particular person sees
Assistants personalise. A logged-in customer with chat history, saved preferences and a phone GPS fix may get an answer we never see. We measure a clean session from your town, which is the closest reproducible thing to a stranger asking cold.
Why a model chose what it chose
No assistant publishes its reasoning or its ranking, and none of them has one in the sense a search engine does. We can show you what came back and what is mechanically wrong on your site. We cannot show you a causal line between the two, and anybody drawing one for you is guessing.
Anything about training data
Whether a model was trained on your site, and what it retained, is not observable from outside. Where an answer names you without searching, that is suggestive. It is not proof, and we label it as suggestive.
Tomorrow
Model versions change without notice and answers drift. A reading is true of the date printed on it and nothing else, which is why every figure in a report carries one.
What we will not claim from this data
That a business is ranked in an AI answer, that we can guarantee it will be named, or that there is a known number of weeks before anything changes. There is no independent evidence that any intervention reliably moves AI citation on a ninety-day cycle, our own included. We will publish our results as we get them, including the ones that go nowhere.
We also do not use this data to write “ChatGPT recommends your competitor instead of you” as though it were a statement about the product OpenAI sells. What we observed was one model, on one date, with search on, from one town, three times — and the report says exactly that.