Measurement, not scoring

Ask it again.

Here is a question a real shopper types. Read the answer. Then press the button and watch the same assistant answer the same question differently, naming different brands and citing different pages. Nothing changed except the roll of the dice.

what's the best sunscreen for kids with sensitive skin?

Press ask to send the question.

your brandcompetitor

Before anything else

Five things to unlearn

Most of the confusion in a first conversation comes from people importing search engine habits into a place where they do not apply. Five corrections do most of the work.

Assistants answer. They do not rank. Your brand is named or it is not, recommended or merely listed, described accurately or not, and given a place to buy or left hanging. Those are the outcomes.

Asking where you rank in ChatGPT asks something the medium cannot answer. Any tool that hands you a rank invented it.

One analysis of 23,387 citations across 240 branded queries found earned media supplying 48 percent of what assistants cite, other third-party commercial content 30 percent, and the brand's own site 23 percent.

So if the answer about your category is assembled from a consumer magazine study, a regulator's report and two retailer product pages, then the lever is what those say. That work is public relations and retail content, not a website migration.

You just watched this happen. Published measurement work puts source overlap between identical same-day queries somewhere between a third and a half.

One screenshot is an anecdote. A measurement is many asks, and every number in the report is a rate over those asks. This is also what makes the exercise cost something, so it is worth establishing early.

A question does not have to be popular to be worth asking. If an assistant has formed a view about a safety concern in your category, that view gets repeated to whoever does ask.

You want to know what it says now, not after a journalist quotes it back to you.

Almost everyone arrives wanting the optimization. Measure first, because without a baseline you cannot attribute any later change to anything you did.

The honest version of this business is that run one tells you where you stand and what is citing your competitors. Run two is where you find out whether the work moved it.

The design choice everything rests on

Your website is the wrong instrument

Nearly every tool in this space inherited its frame from SEO. It crawls your site, scores your pages, and recommends content so the site gets cited. That is a coherent product for a company whose revenue arrives through its own website.

It is the wrong frame for most of the consumer goods market. A sunscreen brand sells through pharmacies and supermarkets. Its site is a brochure with sun-safety tips on it. No amount of on-page work changes the fact that the purchase happens on a shelf, and site traffic is not a proxy for anything the marketing team is trying to move.

Build the measurement around the site and you produce a report full of numbers nobody can act on. So we build it around the category conversation instead, which leads to one rule that governs the whole prompt bank.

A question meant to measure whether your brand comes up on its own must never name your brand, its aliases, its sub-brands, or its own web properties.

Name the brand and the assistant will happily discuss it. The resulting number measures your question, not your standing. This is why most of a good prompt bank never mentions the client.

Where the purchase actually happens

  1. 01

    Asks an assistant what to buy

    Gets a shortlist. Your site is not in it.

  2. 02

    Checks a pharmacy chain's app

    Reads a product page the retailer wrote.

  3. 03

    Reads two reviews

    A parenting forum and a consumer magazine.

  4. 04

    Buys it in the aisle

    Decision already made before arriving.

Times the shopper visited the brand's website: 0

Four outcomes replace traffic

If you cannot count visits, count these instead. Each one is observable in the text of the answer, with no analytics integration and no tag on your site.

01

Mentioned

Does the brand come up at all in a question that never names it?

The floor. Presence in the category conversation.

02

Recommended

Does the assistant suggest it, or merely list it among others?

Being named in a list of eight is not the same as being told to buy it.

03

Named best

Is it singled out rather than grouped?

The strongest form of presence, and the rarest.

04

Given a purchase path

Does the answer tell the person where to buy it?

The closest this gets to purchase intent, and it needs no analytics, because what we record is what the assistant said and not what the person did next.

One caution worth building in before the run rather than after. A mention of your online store is a purchase signal. A mention of your corporate site is awareness. A social account with no shop attached is neither. Average them into one owned-property number and you will report presence as intent.

The system

How a measurement actually runs

Five steps. Click any of them. The parts that carry the judgement are the prompt bank and the reading of the answers, and those are the two the rest of this page is about.

step 1 of 5

Intake

Understand the brand and what the result has to change

Three questions do most of the work. Does the money arrive through your website or through a shelf? What decision would you make differently depending on the result? Who do you compete with, including the brands an assistant treats as interchangeable with yours even if you do not?

That last one catches people out. Assistants routinely group a premium brand with the supermarket own-label sitting next to it. If your competitor list only contains brands your commercial team respects, the share number will be wrong in a way nobody notices.

The minimum to start is a competitor list and your owned domains. Search Console terms, call-centre FAQs and on-site search logs make the questions better aimed. They do not change whether the number is trustworthy.

Writing questions that measure something

Two questions can look different and measure the same thing. What makes them genuinely different is the axis they move along, so the bank is built by deciding the axes first and then filling them.

The axes that make questions different

Who is asking
"sunscreen for a baby under six months" against "sunscreen for an adult with rosacea"
The situation
"a day at the beach" against "walking to work in the city"
Where they are in deciding
"which sunscreen is best" against "how do I stop my kids burning at the beach"
Category or product
"mineral versus chemical sunscreen" against "best SPF 50 for daily use"
Comparison
"which sunscreen brands do dermatologists trust"
Risk and safety
"is oxybenzone safe for children"

Five questions minimum per reported slice

If you are going to put a percentage in the report for a slice, that slice needs at least five questions underneath it. Not five per axis. Five per value of the axis. Decide at intake which slices you will report, because that decision sets the size of the bank rather than the other way round.

Build the branded twin

For key questions, write the unbranded version and the branded version. "Where can I buy mineral sunscreen" measures whether you get found. "Where can I buy Solaria" measures what happens once someone has already chosen you. The second one is where you find out the assistant is sending your customers to a retailer that stopped stocking you.

The surfaces are not interchangeable

Three different things get grouped under the phrase AI answers. They behave differently enough that mixing them into a single score measures nothing in particular, so we report them apart.

never retrieves

A model call

No retrieval. The answer comes from what the model absorbed in training.

Your standing in the model's prior. This is the closest thing that exists to unaided brand awareness in this medium, and it is underrated.

sometimes retrieves

A model call with search

The assistant decides whether to look something up, then writes from what it found plus what it knew.

What most consumers actually see. Also the messiest to measure, because whether retrieval happened at all varies between identical asks.

always retrieves

A search product

Google AI Mode and AI Overviews. A search engine that answers instead of listing.

The bridge between SEO and this work. If anyone in the room still owns an SEO budget, this is the surface that connects their world to yours.

Assistants differ enormously in how much they cite. Some name brands constantly while linking to almost nothing, which is why a surface exposing no sources is recorded as its own state rather than as a zero.

What gets counted

The same things, the same way, in every answer, in every run. Consistency is what makes run two comparable to run one.

  • Brand mentionsEvery appearance of the brand, its aliases, its sub-brands and its misspellings.
  • Competitor mentionsThe same, for every brand in the competitor set, which is what makes a share figure possible.
  • RecommendationsSeparated from mere mentions, because being listed and being recommended are different outcomes.
  • SourcesEvery site cited or named, deduplicated and sorted by who can influence it.
  • Refusals and non-answersRecorded as their own state. An assistant declining to recommend is information, not a missing value.

And whatever else the brand needs

The counters above ship with every engagement. These get added when the brand has a reason for them, which is where most of the personalisation lives.

  • Sentiment toward the brand and toward the category
  • Alerts when an assistant repeats a specific criticism or an outdated study
  • Which retailers and channels the answer names as places to buy
  • Whether a claim you are legally allowed to make is being made for you
  • Whether a sub-brand is being credited to the parent, or the reverse

What lands on your desk

A report you can argue with

Below is a sample run for Solaria, an invented sunscreen brand sold through Mexican pharmacies. Three views of the same 540 answers. Open all three and you will end up diagnosing the brand yourself, which is the point: a score tells you a number, and this tells you what to go and do on Monday.

Sample run60Questions3Assistants3Repetitions540Answers read396Unbranded subset

Rates, and what each one is a rate over

Every figure is a share of the answers it was measured in, and the denominator is printed next to it. A percentage with no population behind it is a decoration.

The four outcomes

n = 396 answers to unbranded questions

Mentioned26%

103 of 396 answers name the brand unprompted

Recommended11%

44 answers actively suggest it

Named best3%

12 answers single it out

Given a purchase path7%

28 answers say where to buy it

Share of category mentions

n = Every brand named across the same 396 answers

Bloqsol34%
Dermalux27%
Solaria18%
Vera Sun12%
Supermarket own-label9%

Mention rate by surface

n = 132 unbranded answers per surface

Model call, no retrieval9%

What the model already believed

Model call with search22%

What most consumers see

Google AI Mode31%

Search that answers instead of listing

Solaria is mentioned in about a quarter of category answers and recommended in one in nine. The gap between those two is the whole opportunity: the assistants know the brand exists and mostly decline to suggest it.

Invented brand, invented competitors, invented figures. The structure is real; the numbers are not anyone's.

The one-pager

Everything above compresses onto one page a marketing director can carry into a meeting. What was measured, the four rates, the sources that matter, and the handful of things worth doing. Have a look at it before you decide whether any of this is useful to you.

Size it yourself

Three dials, and you control all of them. The observation count is fixed before the first question is sent, which means the scope conversation happens before the money is spent rather than after.

60
3
3

60 × 3 × 3

Answers to read

540

Roughly 9 hours of human reading, at about a minute an answer.

What makes this expensive is reading the answers, not asking the questions. The API cost of asking is close to nothing. Every sizing decision above is really a decision about how much reading somebody has signed up for, which is why we make you look at the number before you commit to it.

Why anyone runs it twice

The loop is the whole point

Run one gives you a baseline and a list of what is citing your competitors. Then the work happens, which is usually editorial, retail content and public relations rather than a website project.

Then you ask the same frozen bank again. Same questions, same assistants, same number of repetitions. Any movement is attributable, because the instrument did not change.

This is the part that turns a one-off study into something worth keeping, and it is also the part that keeps us honest. If nothing moved, the report says nothing moved.

The bank is frozen after run one. Add a question later and you have two banks, not one trend, so new questions go into a separate set that starts its own baseline.

Solaria recommended, unbranded questions

illustrative figures

Run 1 · baseline11%
Run 2 · after the work26%

+15 points. The bank did not change between the two runs, so the movement belongs to the work rather than to a different set of questions.

What is different here

Three things you will not get elsewhere

Nothing to buy, nothing to learn

There is no dashboard, no seat licence and no tool your team has to adopt. A specialist runs the measurement and you receive the analysis and the report. Most brand teams already have more software than they use, and the last thing a quarterly measurement needs is a login somebody forgets.

Built for brands that do not sell online

Every tool in this category crawls your website and scores your pages. If your product is bought off a shelf, that answers a question you did not ask. This is built the other way round, from the category conversation inward, which is why it works for a packaged food brand and works just as well for a software company that does sell online.

You learn the method, not just the number

Part of the engagement is teaching your team how assistants build answers, why the bank is written the way it is, and how to read a rate. The aim is that you can defend these numbers in a meeting we are not in. A number nobody on your side can explain does not survive its first challenge.

What this does not do

Everyone else in this category sells an instant score. Here is where the boundaries of an honest measurement actually sit, so you can decide before you spend anything.

It will not give you a rank
There is no position one to report. Anything that gives you one made it up.
It will not attribute sales
We record what the assistant said. What the person did next happens outside anything we can see. The purchase-path outcome is the closest honest proxy, and it is a proxy.
It cannot promise the number will move
Most of what an assistant cites was written by somebody else. We can tell you which sources are driving the answer and who can influence each one. We cannot commit to a result that depends on third parties publishing differently.
It does not fact-check the assistant for you
We can record that an assistant said something about your product. Judging whether that statement is true needs an approved source of truth that somebody on your side maintains, and most brands do not have one. If accuracy is your worry, say so at intake, because building that reference is a separate piece of work.
One run tells you where you stand, not what to do
The action plan comes from the source list and the gaps, and it gets sharper with a second run. Anyone promising a fix from a single baseline is selling you the wrong thing.

Start with the question, not the tool

The useful first conversation is about what decision the result should change, and whether your category is one where an assistant already has opinions. That takes about half an hour, and you will know by the end whether measuring is worth it for you.

Talk it through