Blog · Measurement · 6 min read

Why ChatGPT gives different answers to the same question, and what a margin of error fixes

Why does ChatGPT provide different answers to the same question? Sampling and retrieval create variance, while margins of error show whether a change is real.

“Surely this is just SEO” is a reasonable objection. If ChatGPT gives different answers to the same question, the obvious explanation is that your pages moved, a competitor published something better, or the engine changed what it can retrieve. Those things matter. They are not the whole explanation.

The question is simple: why does ChatGPT provide different answers to the same question? A response is not a fixed ranking position, but an observation from a process with several moving parts. The question wording, conversation context and retrieved pages can change what the system selects. A result can therefore look like a visibility win, a visibility loss, and a return to normal within a short period.

Variance

A single answer tells you what happened once. It does not tell you the underlying rate at which ChatGPT names your product or cites your page. That distinction matters when you compare a page change with an earlier result. A result that appears in one answer and disappears in the next may reflect retrieval variance rather than a real change in your site.

Our sample makes the problem visible at a larger scale. We have recorded 64,667 citations across 9,791 AI answers, from 77 scans of 44 sites in 48 categories. The scans asked 507 distinct buyer questions and covered 9,192 distinct domains and 21,465 distinct pages, according to the CiteGraph research data. That is enough evidence to measure patterns. It is not evidence that every individual answer is stable.

The engine matters too. ChatGPT named the scanned product in 6.1% of 2,434 sampled answers, while Perplexity did so in 10.5% of 2,440 answers, Gemini in 9.4% of 2,437, and Claude in 10.7% of 2,120. ChatGPT also cited an average of 3.5 sources per answer, compared with 5.5 for Perplexity, 4.2 for Gemini and 4.9 for Claude. These figures describe different samples and engines, so they are not a league table for your business. They show why a result from one assistant cannot stand in for AI visibility as a whole.

Sources cited per answer, by enginePerplexity: 5.5; Gemini: 4.2; ChatGPT: 3.5; Claude: 4.9 SOURCES CITED PER ANSWER, BY ENGINE Perplexity 5.5 Gemini 4.2 ChatGPT 3.5 Claude 4.9 Source: CiteGraph data
Average sources cited per answer varied by engine in the sampled answers.

Sampling creates the first kind of variance. The question wording, the conversation context and the pages retrieved can differ even when the prompt looks identical to you. Retrieval creates the second: the engine selects a small working set from a much larger web, then decides which sources support the answer. Model updates create the third, because the system that interprets those sources can change without your page changing at all.

This does not mean ChatGPT is broken. A reader who asks ChatGPT the same arithmetic question twice and gets different working is describing something different from a citation disappearing. Arithmetic has a correct result under fixed assumptions, while product recommendations and source selection depend on context, retrieval and generation. ChatGPT's single reply box gives no signal that these are different kinds of failure.

Retrieval

The useful measurement separates product naming, domain citation and exact-page citation. A domain can remain part of the retrieval pool while the page shown in an answer changes. A page can also be selected once because it matches the wording particularly well, then fail to appear when another relevant source takes its place. These measures answer different business questions, so combining them into one pass or fail number hides the movement that matters.

Our project on source churn shows the distinction. Of the 40 domains cited three or more times in the first scan of our own site, 27 were still cited 25 days later. Of 179 exact pages cited in that first scan, 14 were cited again 25 days later. Looking across every cited domain, 56 of 102 came back. These are illustrative results, not statistically tight estimates: with samples of 40, 179 and 102, the intervals around the observed rates are wide, especially for the page result. On this sample, domain-level citation looks more persistent than page-level citation, though the page count of 179 makes that comparison noisier than the domain counts.

What survived 25 days68% Heavily cited domains; 55% All cited domains; 8% Exact pages WHAT SURVIVED 25 DAYS 68% HEAVILY CITED DOMAINS 55% ALL CITED DOMAINS 8% EXACT PAGES Source: CiteGraph data
Domains were more persistent than exact pages across scans 25 days apart.

The long tail is mostly sampling, not change, in this project. That does not make page-level movement irrelevant. It means you need to decide which level matters for the business question before interpreting a different answer. A page rewrite may improve the chance that your domain is cited while the exact URL varies. A clearer product statement may increase naming without increasing citations.

Margins

A margin of error turns a noisy count into a decision rule. Suppose you run the same tracked question repeatedly and observe an outcome in a share of runs. Call that observed share p and the number of runs n. For a simple proportion, the uncertainty shrinks in proportion to the square root of n. In practical terms, the runs you need scale like this:

new runs = current runs × (current margin / target margin)²

Choose the confidence method, calculate the interval around the observed rate, then compare the intervals for the two periods. If the intervals overlap heavily, describe the movement as directional. If the interval around the new result sits clear of the old baseline, you have stronger evidence that the change is real. The exact number of runs depends on the baseline rate and the margin you are willing to accept, neither of which is supplied by a single prompt.

A better report stores the prompt, engine, date, answer, cited domains, cited URLs and outcome definition. It repeats the observation, calculates the rate and prints the margin. Our measurement method guide explains the mechanics, while the CiteGraph methodology sets out how we measure citations and source persistence. This gives you a record you can inspect when a page change and an engine change happen close together.

The formula also gives you a way to plan a test. If you changed a page and want to know whether citation rate moved, keep the question set and engine treatment as consistent as possible. Run enough observations to make the target margin useful, rather than choosing a round count because it feels official. A larger sample does not remove retrieval or model variance. It makes your estimate less easily fooled by it.

A short test you can run in under an hour

  1. Choose a fixed set of buyer questions that already matter to your business.
  2. Record the engine, wording, date, named products, cited domains and exact cited pages.
  3. Repeat the same questions without changing the outcome definition.
  4. Calculate the share for each outcome and attach a margin of error.
  5. Compare the result with the earlier baseline, then label it as stable, directional or materially different.

Do not change the prompt halfway through and call the result a before and after. That tests question wording as well as your page. Do not combine naming and citation into one measure. They can move independently.

Limits

Here is what we cannot tell you. Our data measures samples of AI answers, not every ChatGPT response, every user context or every model state. We can describe the behaviour in the 77 scans and the 44 sites included in the cited sample, but we cannot turn those observations into a guaranteed probability for your next prompt. We also cannot isolate the exact contribution of sampling, retrieval and model updates from the aggregate churn figures alone.

That limit is important for a founder deciding whether to publish, rewrite or wait. A result that changes after 25 days may reflect a model update, a retrieval change, a different sample, or some combination. The data does not identify the cause by itself. We measure samples; we print margins; we do not know what we have not measured.

The sensible response is not to ignore SEO. Crawl access, clear product facts and useful comparison pages still affect whether an engine has material it can use. The argument against our position is that a well-run SEO programme should already handle most of this: stable pages, strong sources and consistent measurement should reduce noise. That is partly right. Better pages can improve the underlying chance of selection, but they cannot make a sampled generative system deterministic. Treating variance as an excuse for weak content, or a single answer as proof of success, is a mistake either way.

For a working check, start with the free AI visibility checker to inspect the answers, named products and sources attached to your questions. Use the weekly research on who owns AI answers when you need broader context, not as a substitute for your own baseline. The first tells you what to record for the procedure above, while the second helps you see whether a pattern persists beyond your own site. The goal is a repeatable record of what the engine did, with enough observations to tell a page change from ordinary movement.

ChatGPT gives different answers because each one is a sample, not a fixed result.

Questions people ask

How many ChatGPT runs do I need before I can trust a change?+

There is no fixed run count that works for every question. Use the formula `new runs = current runs × (current margin / target margin)²`, then compare the resulting interval with your earlier baseline.

Can a page lose a ChatGPT citation without losing visibility?+

Yes. The domain can remain in the retrieval pool while another page replaces the exact URL in the answer. Measure domain citation, page citation and product naming separately.

If two people ask ChatGPT the same thing, will they see my brand named the same way?+

Not necessarily. Context, retrieval and sampling can differ between interactions, so one response cannot establish the normal naming rate for your brand.

Should I change my SEO strategy when ChatGPT answers are inconsistent?+

Keep improving access, clear product facts and pages that answer buyer questions. Measure the resulting citation and naming rates over repeated runs instead of reacting to one answer.

Can I trust one ChatGPT answer when a buyer is deciding what to buy?+

Treat it as one observation, not as a stable visibility rate. Check repeated answers, the cited sources and the margin around the outcome before drawing a business conclusion.

Sources
  1. CiteGraph scan data, read live citegraph.app/research
  2. CiteGraph's own project, scans 25 days apart citegraph.app/methodology

Cite this: CiteGraph, “Why ChatGPT gives different answers to the same question, and what a margin of error fixes”, 21 Sept 2026, https://www.citegraph.app/blog/why-chatgpt-gives-different-answers-to-the-same-question

Share: X · LinkedIn · Email

Recommended reading
Measurement · 6 min

How to track brand mentions in ChatGPT

Learn how to track brand mentions in ChatGPT by hand, with an API script, or using a monitoring tool, while keeping the sample stable enough to trust.

21 Sept 2026
Getting recommended · 5 min

How to rank in ChatGPT answers, not search results

Learn how to improve your product’s position in ChatGPT answers by analysing cited pages, source order, page types and answer stability.

11 Sept 2026
Measurement · 6 min

Cross-engine AI citation tracking: why one engine is not enough

Cross-engine AI citation tracking shows why one assistant misses source differences, engine gaps and the pages buyers actually see.

21 Sept 2026