Measurement, not scoring
Ask it again.
Here is a question a real shopper types. Read the answer. Then press the button and watch the same assistant answer the same question differently, naming different brands and citing different pages. Nothing changed except the roll of the dice.
what's the best sunscreen for kids with sensitive skin?
readyPress ask to send the question.
Before anything else
Five things to unlearn
Most of the confusion in a first conversation comes from people importing search engine habits into a place where they do not apply. Five corrections do most of the work.
Assistants answer. They do not rank. Your brand is named or it is not, recommended or merely listed, described accurately or not, and given a place to buy or left hanging. Those are the outcomes.
Asking where you rank in ChatGPT asks something the medium cannot answer. Any tool that hands you a rank invented it.
One analysis of 23,387 citations across 240 branded queries found earned media supplying 48 percent of what assistants cite, other third-party commercial content 30 percent, and the brand's own site 23 percent.
So if the answer about your category is assembled from a consumer magazine study, a regulator's report and two retailer product pages, then the lever is what those say. That work is public relations and retail content, not a website migration.
You just watched this happen. Published measurement work puts source overlap between identical same-day queries somewhere between a third and a half.
One screenshot is an anecdote. A measurement is many asks, and every number in the report is a rate over those asks. This is also what makes the exercise cost something, so it is worth establishing early.
A question does not have to be popular to be worth asking. If an assistant has formed a view about a safety concern in your category, that view gets repeated to whoever does ask.
You want to know what it says now, not after a journalist quotes it back to you.
Almost everyone arrives wanting the optimization. Measure first, because without a baseline you cannot attribute any later change to anything you did.
The honest version of this business is that run one tells you where you stand and what is citing your competitors. Run two is where you find out whether the work moved it.
The design choice everything rests on
Your website is the wrong instrument
Nearly every tool in this space inherited its frame from SEO. It crawls your site, scores your pages, and recommends content so the site gets cited. That is a coherent product for a company whose revenue arrives through its own website.
It is the wrong frame for most of the consumer goods market. A sunscreen brand sells through pharmacies and supermarkets. Its site is a brochure with sun-safety tips on it. No amount of on-page work changes the fact that the purchase happens on a shelf, and site traffic is not a proxy for anything the marketing team is trying to move.
Build the measurement around the site and you produce a report full of numbers nobody can act on. So we build it around the category conversation instead, which leads to one rule that governs the whole prompt bank.
A question meant to measure whether your brand comes up on its own must never name your brand, its aliases, its sub-brands, or its own web properties.
Name the brand and the assistant will happily discuss it. The resulting number measures your question, not your standing. This is why most of a good prompt bank never mentions the client.
Where the purchase actually happens
- 01
Asks an assistant what to buy
Gets a shortlist. Your site is not in it.
- 02
Checks a pharmacy chain's app
Reads a product page the retailer wrote.
- 03
Reads two reviews
A parenting forum and a consumer magazine.
- 04
Buys it in the aisle
Decision already made before arriving.
Times the shopper visited the brand's website: 0
Four outcomes replace traffic
If you cannot count visits, count these instead. Each one is observable in the text of the answer, with no analytics integration and no tag on your site.
Mentioned
Does the brand come up at all in a question that never names it?
The floor. Presence in the category conversation.
Recommended
Does the assistant suggest it, or merely list it among others?
Being named in a list of eight is not the same as being told to buy it.
Named best
Is it singled out rather than grouped?
The strongest form of presence, and the rarest.
Given a purchase path
Does the answer tell the person where to buy it?
The closest this gets to purchase intent, and it needs no analytics, because what we record is what the assistant said and not what the person did next.
One caution worth building in before the run rather than after. A mention of your online store is a purchase signal. A mention of your corporate site is awareness. A social account with no shop attached is neither. Average them into one owned-property number and you will report presence as intent.
The system
How a measurement actually runs
Five steps. Click any of them. The parts that carry the judgement are the prompt bank and the reading of the answers, and those are the two the rest of this page is about.
step 1 of 5
Intake
Understand the brand and what the result has to change
Writing questions that measure something
Two questions can look different and measure the same thing. What makes them genuinely different is the axis they move along, so the bank is built by deciding the axes first and then filling them.
The axes that make questions different
- Who is asking
- "sunscreen for a baby under six months" against "sunscreen for an adult with rosacea"
- The situation
- "a day at the beach" against "walking to work in the city"
- Where they are in deciding
- "which sunscreen is best" against "how do I stop my kids burning at the beach"
- Category or product
- "mineral versus chemical sunscreen" against "best SPF 50 for daily use"
- Comparison
- "which sunscreen brands do dermatologists trust"
- Risk and safety
- "is oxybenzone safe for children"
Five questions minimum per reported slice
If you are going to put a percentage in the report for a slice, that slice needs at least five questions underneath it. Not five per axis. Five per value of the axis. Decide at intake which slices you will report, because that decision sets the size of the bank rather than the other way round.
Build the branded twin
For key questions, write the unbranded version and the branded version. "Where can I buy mineral sunscreen" measures whether you get found. "Where can I buy Solaria" measures what happens once someone has already chosen you. The second one is where you find out the assistant is sending your customers to a retailer that stopped stocking you.
The surfaces are not interchangeable
Three different things get grouped under the phrase AI answers. They behave differently enough that mixing them into a single score measures nothing in particular, so we report them apart.
A model call
No retrieval. The answer comes from what the model absorbed in training.
Your standing in the model's prior. This is the closest thing that exists to unaided brand awareness in this medium, and it is underrated.
A model call with search
The assistant decides whether to look something up, then writes from what it found plus what it knew.
What most consumers actually see. Also the messiest to measure, because whether retrieval happened at all varies between identical asks.
A search product
Google AI Mode and AI Overviews. A search engine that answers instead of listing.
The bridge between SEO and this work. If anyone in the room still owns an SEO budget, this is the surface that connects their world to yours.
Assistants differ enormously in how much they cite. Some name brands constantly while linking to almost nothing, which is why a surface exposing no sources is recorded as its own state rather than as a zero.
What gets counted
The same things, the same way, in every answer, in every run. Consistency is what makes run two comparable to run one.
- Brand mentionsEvery appearance of the brand, its aliases, its sub-brands and its misspellings.
- Competitor mentionsThe same, for every brand in the competitor set, which is what makes a share figure possible.
- RecommendationsSeparated from mere mentions, because being listed and being recommended are different outcomes.
- SourcesEvery site cited or named, deduplicated and sorted by who can influence it.
- Refusals and non-answersRecorded as their own state. An assistant declining to recommend is information, not a missing value.
And whatever else the brand needs
The counters above ship with every engagement. These get added when the brand has a reason for them, which is where most of the personalisation lives.
- Sentiment toward the brand and toward the category
- Alerts when an assistant repeats a specific criticism or an outdated study
- Which retailers and channels the answer names as places to buy
- Whether a claim you are legally allowed to make is being made for you
- Whether a sub-brand is being credited to the parent, or the reverse
What lands on your desk
A report you can argue with
Below is a sample run for Solaria, an invented sunscreen brand sold through Mexican pharmacies. Three views of the same 540 answers. Open all three and you will end up diagnosing the brand yourself, which is the point: a score tells you a number, and this tells you what to go and do on Monday.
Rates, and what each one is a rate over
Every figure is a share of the answers it was measured in, and the denominator is printed next to it. A percentage with no population behind it is a decoration.
The four outcomes
n = 396 answers to unbranded questions
103 of 396 answers name the brand unprompted
44 answers actively suggest it
12 answers single it out
28 answers say where to buy it
Share of category mentions
n = Every brand named across the same 396 answers
Mention rate by surface
n = 132 unbranded answers per surface
What the model already believed
What most consumers see
Search that answers instead of listing
Solaria is mentioned in about a quarter of category answers and recommended in one in nine. The gap between those two is the whole opportunity: the assistants know the brand exists and mostly decline to suggest it.
Invented brand, invented competitors, invented figures. The structure is real; the numbers are not anyone's.
The one-pager
Everything above compresses onto one page a marketing director can carry into a meeting. What was measured, the four rates, the sources that matter, and the handful of things worth doing. Have a look at it before you decide whether any of this is useful to you.
Size it yourself
Three dials, and you control all of them. The observation count is fixed before the first question is sent, which means the scope conversation happens before the money is spent rather than after.
60 × 3 × 3
Answers to read
540
Roughly 9 hours of human reading, at about a minute an answer.
What makes this expensive is reading the answers, not asking the questions. The API cost of asking is close to nothing. Every sizing decision above is really a decision about how much reading somebody has signed up for, which is why we make you look at the number before you commit to it.
Why anyone runs it twice
The loop is the whole point
Run one gives you a baseline and a list of what is citing your competitors. Then the work happens, which is usually editorial, retail content and public relations rather than a website project.
Then you ask the same frozen bank again. Same questions, same assistants, same number of repetitions. Any movement is attributable, because the instrument did not change.
This is the part that turns a one-off study into something worth keeping, and it is also the part that keeps us honest. If nothing moved, the report says nothing moved.
The bank is frozen after run one. Add a question later and you have two banks, not one trend, so new questions go into a separate set that starts its own baseline.
illustrative figures
+15 points. The bank did not change between the two runs, so the movement belongs to the work rather than to a different set of questions.
What is different here
Three things you will not get elsewhere
Nothing to buy, nothing to learn
There is no dashboard, no seat licence and no tool your team has to adopt. A specialist runs the measurement and you receive the analysis and the report. Most brand teams already have more software than they use, and the last thing a quarterly measurement needs is a login somebody forgets.
Built for brands that do not sell online
Every tool in this category crawls your website and scores your pages. If your product is bought off a shelf, that answers a question you did not ask. This is built the other way round, from the category conversation inward, which is why it works for a packaged food brand and works just as well for a software company that does sell online.
You learn the method, not just the number
Part of the engagement is teaching your team how assistants build answers, why the bank is written the way it is, and how to read a rate. The aim is that you can defend these numbers in a meeting we are not in. A number nobody on your side can explain does not survive its first challenge.
What this does not do
Everyone else in this category sells an instant score. Here is where the boundaries of an honest measurement actually sit, so you can decide before you spend anything.
- It will not give you a rank
- There is no position one to report. Anything that gives you one made it up.
- It will not attribute sales
- We record what the assistant said. What the person did next happens outside anything we can see. The purchase-path outcome is the closest honest proxy, and it is a proxy.
- It cannot promise the number will move
- Most of what an assistant cites was written by somebody else. We can tell you which sources are driving the answer and who can influence each one. We cannot commit to a result that depends on third parties publishing differently.
- It does not fact-check the assistant for you
- We can record that an assistant said something about your product. Judging whether that statement is true needs an approved source of truth that somebody on your side maintains, and most brands do not have one. If accuracy is your worry, say so at intake, because building that reference is a separate piece of work.
- One run tells you where you stand, not what to do
- The action plan comes from the source list and the gaps, and it gets sharper with a second run. Anyone promising a fix from a single baseline is selling you the wrong thing.
Start with the question, not the tool
The useful first conversation is about what decision the result should change, and whether your category is one where an assistant already has opinions. That takes about half an hour, and you will know by the end whether measuring is worth it for you.
Talk it through