Method · reproducible with a spreadsheet and API access

How to measure brand citations in ChatGPT and other AI assistants

Most advice on AI visibility is a screenshot with a conclusion attached. This is the method behind a number you could defend: what to ask, how many times, what to count, how to score a miss, and how wide the error bars are. Worked examples come from 5,639 real answers.

By CiteGraph · Updated 2026-09-07 · Prices checked on the vendors’ own pages

CiteGraph is a tool that does this automatically, and this page describes the method it uses. Nothing here needs the tool. A spreadsheet, API access to the engines and an afternoon will reproduce a single measurement; the tool exists because doing it weekly gets old.

1. Prompt design: ask what buyers ask

The measurement is only as good as the questions. Keywords are not questions; nobody types “AI SEO tool” into an assistant. Buyers type situations: “I run a two-person agency, what should I use to see whether ChatGPT recommends my clients?” Write ten of those for one product, and cover the shapes buyers actually use:

Read the product’s own site first so the questions use its buyer’s vocabulary and its real price band. Then fix the set. If the questions change between measurements, you have measured two different things and cannot compare them.

2. Repeated sampling: one answer is a screenshot

Ask each question of each engine several times, through the engine’s API with web search on, so the retrieval matches what a buyer triggers. Record every answer verbatim, including failures. Three runs per engine is the floor; five tightens the interval by about a third. Exclude failed calls from every denominator rather than counting them as misses.

Why it matters, from our own data: across 5,639 answers on 34 products, the scanned product was named in 4.3% of ChatGPT-family answers, 5.7% on Claude, 6.2% on Gemini and 6.4% on Perplexity. At those rates, a single run of ten questions will show zero mentions most of the time and one or two on a lucky day, and the difference between those is nothing.

3. Citation classification: count two things, not one

For every answer, record separately whether the product was named in the prose (strip URLs first, so a link never inflates it) and whether a page on its own domain was cited as a source. Then classify every cited page by kind: listicle, review platform, community thread, directory, news, video, the brand’s own site, a competitor’s site, documentation, other. Unwrap redirect links to the true host before classifying, or Google’s grounding wrapper will show up as a source.

The classification is where the actionable finding lives. Across every scan we have run, reddit.com is cited in 100% of scans and answers 7.3 of 10 buyer questions on average; youtube.com is cited in 75% of scans, linkedin.com in 73%, medium.com in 67% and g2.com in 65%. Engines build recommendations from those pages far more than from any brand’s homepage, and a measurement that does not list them cannot tell you where to go.

4. Missed-opportunity scoring: which loss to fix first

A question where the product was never named is a blind spot. For each one, record what was named instead, which pages the answer came from, and the rate the product would reach overall if it won that question in every run. That last number orders the work: winning a question you lose on all four engines moves the headline rate by a full tenth; winning one you already half-hold moves it far less.

Then read the answers. The reason the engine chose someone else is usually visible in the text: a price the product never states, a use case its site never mentions, a comparison page it does not have. The fix for a blind spot is nearly always a sentence the product’s own pages, or the pages the engine trusts, do not yet contain.

5. Confidence limits: how wide the error bars really are

Every rate is a proportion from a sample, so give it an interval rather than a plain percentage; Wilson at 95% is the one we use on the public boards, and it behaves at small counts. Two worked examples at the rates you will actually see:

SampleNamedRate95% intervalReads as
10 questions × 3 runs × 4 engines = 120 answers65.0%1.5% to 11.3%5% ± about 5 points
10 × 5 × 4 = 200 answers105.0%2.3% to 9.4%5% ± about 3.5 points
1 question × 1 run × 1 engine = 1 answer00%0% to 79%a screenshot

The rule for change: report movement only when the new interval and the old one do not overlap. A product that goes from 5% to 8% on 120 answers has not moved; it has sampled. This is unglamorous and it is the difference between a measurement and a mood.

6. Record it in a shape others can check

Keep the questions, the engines and how they were queried, every answer with its named and cited flags and its classified sources, and the summary with intervals. We publish a JSON schema for exactly that record, free to use, so measurements from different people or tools can be laid side by side:

citegraph.app/schemas/citation-measurement.schema.json

It covers the product, the fixed question set with intent, the engines with the model and whether web search was on, one record per question × engine × run with the verbatim text, and a summary with named and cited rates, their margins, per-engine breakdowns and blind spots with the if-won rate. A measurement that fits the schema is one somebody else can re-run.

What this method cannot tell you

Whether an engine’s claim about a product is true; only that it made it, attributed to that engine and that run. Whether a buyer acted on the answer. And anything about engines queried through a scraped chat session rather than an API, because those sessions carry history and personalisation the method cannot control for. State those limits with the numbers, and the numbers stay honest.

Common questions

How do I check whether ChatGPT mentions my brand?+

Ask it the questions your buyers ask, in their words, several times each, and record every answer verbatim. Count the answers that name your product and the answers that cite a page on your site, separately, and put a margin of error on both. One answer proves nothing; the same question asked twice returns different products.

How many times should each question be asked?+

At least three per engine, five if you can afford it. With ten questions on four engines, three runs gives 120 answers and a margin of about five points on a rate near 5%; five runs gives 200 answers and about three and a half points.

What is the difference between named and cited?+

Named: your product appears in the answer's text. Cited: a page on your own domain is among the sources the answer used. They diverge constantly. Across 5,639 answers we have measured, the scanned product was named in about 5% of answers and cited far less often, because engines learn about products from listicles, forums and review sites rather than from the brand's site.

How do I know a change is real?+

Put a 95% interval on each rate and call a change movement only when the two intervals do not overlap. Everything smaller is within sampling noise and should be reported as no change.

Is there a standard format for recording this?+

Not an industry one. We publish a JSON schema for a measurement record at citegraph.app/schemas/citation-measurement.schema.json so measurements can be shared and compared in the same shape. It is free to use.

See which pages answer your category — and what they say about you.

Check your site →