Error bars on every number
Ask an AI engine the same question twice and you'll get two different answers. Any tool reporting visibility from one run per prompt is reporting an anecdote with a percent sign. CiteGraph runs every question 3–5 times on each engine and prints the margin of error next to every rate — which changes what you do with the number.
How it works
Repeated runs, real distributions
Each buyer question runs 3–5 times per engine. The rate is computed across all cells; the margin comes from the binomial spread, with the rule of three at the zero boundary.
Margins printed, not footnoted
±9 sits next to 23% wherever 23% appears — dossier, dashboard, digest. When two brands overlap inside their margins, the UI says tied.
Movement must beat the margin
Week-over-week changes are only announced when they clear the printed margin. Everything else is honestly labeled as within noise.
In depth
Why competitors don't do this
Printing a margin means admitting your daily delta chart is mostly noise — which is hard when daily deltas are the product. Repeated runs also cost real API money per scan. We pay it because a number you can't trust is worth exactly nothing.
Zero is measured too
Being named in 0 of 120 answers doesn't mean exactly 0% — the rule of three puts the honest ceiling around 3%. Even our zeros come with an interval, including the one we published about ourselves.
Failures shrink the sample, visibly
Engine calls fail. Failed cells are excluded from denominators and the shortfall is printed — a rate over 83 of 90 planned answers says so on the page.
What we won't do: report a single-run rate as truth, or animate a daily chart whose wiggles are sampling noise. If that makes our graphs calmer than a rival's, the calm is the accuracy.
Common questions
Why do AI visibility numbers need error bars?+
Because engines are probabilistic: retrieval, model sampling, and answer length all vary between runs. A brand at "23%" from one run per prompt could genuinely be anywhere from 10% to 40%. The margin is the difference between measurement and vibes.
How is the margin calculated?+
Binomial spread over the answer cells for normal cases; Wilson intervals on public leaderboards; the rule of three (3/n) when a count is zero, so certainty is never claimed at the boundary.
Doesn't repeated running cost more?+
Yes — it's most of a scan's API cost, and why there's no free tier. One honest scan beats five free anecdotes.
See it on your own domain — first dossier in about a minute.