How to choose an AI citation tracking tool
Six things a tracker has to get right, each with a test you can run on any vendor before paying. Competitor facts are published prices and plans from their own pages, dated; everything we could not verify is marked as such.
By CiteGraph · Updated 2026-09-07 · Prices checked on the vendors’ own pages
Disclosure: CiteGraph makes one of the tools in this category. The framework below is the one we built the product against, so it favours the things we do. The tests are the defence: they work on us too, and you should run them on us before believing this page.
What you are actually buying
An AI citation tracker answers one question: when a buyer asks an assistant what to use, does it name you, and if not, who and why? Everything else on a vendor’s feature list is in service of that. So the criteria are about whether the number is real, whether it can be acted on, and whether it will still mean the same thing next month.
1. Prompt sampling
The questions decide everything. Keyword-style prompts ("AI SEO tool") measure something no buyer types; buyer-style prompts ("what should a two-person agency use to see whether ChatGPT recommends its clients?") measure the thing you care about.
The test
Ask to see the exact prompts before you pay. If you can't read them, edit them and keep the same set week to week, movement will be a different question, not a different answer.
2. Citation extraction
Engines return sources in different shapes: Perplexity lists them natively, others wrap them in redirect links, some name a brand in prose without linking anything. A tool has to unwrap the redirect to the true host and separate "named in the text" from "cited as a source".
The test
Run one question yourself in the engine, then compare the tool's record of that answer. Does it show the answer verbatim? Does it list the same sources you see? Are named and cited reported as two numbers?
3. Source-level reporting
The pages engines cite are mostly not the brands' own sites. Across every scan we have run, reddit.com is cited in 100% of scans and answers 7 of 10 buyer questions on average. A score without the list of cited pages tells you that you lost and not where.
The test
Ask for the list of cited domains for one question, graded by what each page actually says about your category. If the tool cannot show you the page behind a number, the number cannot be acted on.
4. Recency
Answers change weekly as engines re-crawl and as the pages they cite change. A one-off audit is a photograph; tracking is the film.
The test
Ask how often the same prompt set is re-run, and whether a change is reported only when it clears the margin of error. A tool that reports every wobble as a trend will have you chasing noise.
5. Platform coverage
Buyers use different assistants, and the assistants disagree. In our data, Gemini names a tracked competitor in 63% of answers and Claude in 34%; the same product is rarely treated the same way twice.
The test
Count the engines, then ask how each is queried. Official APIs with web search behave like the products buyers use; scraped chat sessions and cached indexes do not. A tool should be able to tell you which it does.
6. Verification
Every rate is a sample. Without an interval you cannot tell a real move from a coin flip, and without the verbatim answers you cannot check a claim at all.
The test
Ask for the margin of error on any number and the raw answer behind any citation. If either is missing, treat the dashboard as an estimate with unknown precision.
The tools, by the same criteria
Prices and plan facts are from each vendor’s own pricing page, checked 23 August 2026. Where a vendor does not publish something, the cell says so. We did not sign up for competitors’ trials to test the rest, so those cells say “not verified” rather than guess in our own favour.
| Tool | Entry price | Engines | Prompts | Repeated runs | Verbatim answers | Margins |
|---|---|---|---|---|---|---|
| CiteGraph (us) | $29/mo | 4: ChatGPT, Claude, Gemini, Perplexity | 10 per project, editable | 3 per engine (5 on Growth and Scale) | Every answer kept | Printed on every rate (binomial; rule of three at zero) |
| Otterly.AI | $29/mo (Lite) | 4 (per pricing page) | 15 (Lite) | not published | not verified | not published |
| Profound | $99/mo billed yearly | not verified | not published | not published | not verified | not published |
| Peec AI | not published | not verified | not published | not published | not verified | not published |
| CrowdReply | $99/mo (Starter) | choose 2 (Starter) | 20 (Starter) | not published | not verified | not published |
Sources: Otterly.AI otterly.ai/pricing/; Profound www.tryprofound.com/pricing; Peec AI peec.ai/pricing; CrowdReply crowdreply.io/pricing; CiteGraph www.citegraph.app/pricing;
A checklist to take into any trial
Before the trial ends, you should be able to tick every one of these. Any you cannot tick is a thing the tool will not tell you later either.
- I have read the exact prompts and they are questions a buyer would type.
- I can open any number and see the verbatim answers behind it.
- Named and cited are separate figures.
- Every rate shows a margin of error, and the tool says how many answers it came from.
- I can see the list of cited pages for a lost question, not only a score.
- I know which engines are queried and how (API with web search, or something else).
- The same prompt set re-runs on a schedule, and a change is only called a change when it clears the margin.
- I know what the tool wants me to do next, in order, and it names the page to go after.
Where CiteGraph stands on its own criteria
Ten prompts written from your site and editable before anything runs; four engines through official APIs with web search; three runs per engine on Starter and five on Growth and Scale; every answer kept verbatim; named and cited reported separately; a margin printed on every rate (a binomial standard error, the rule of three at zero) and movement only called when it clears it; the cited pages for every question, graded; and an action plan in order. Where we are weak: we are new, so we appear in none of the roundups these answers are built from, and our own product is named in 0% of answers in its category. The methodology page has the rest, including what we deliberately do not do.
Common questions
What is an AI citation tracking tool?+
Software that asks AI assistants such as ChatGPT, Claude, Gemini and Perplexity the questions your buyers ask, records which products the answers name and which web pages they cite, and reports how that changes over time. The useful ones also tell you what to change.
What is the difference between being named and being cited?+
Named means your product appears in the answer's text. Cited means a page on your own site was one of the sources the engine used. A product can be recommended constantly while its own site is never cited, because the engine learned about it from listicles and forums. A tool that counts only one of the two is telling you half the story.
Why does repeated sampling matter?+
The same question asked twice returns different products. One answer is a screenshot, not a measurement. A tool should ask each question several times per engine and publish a margin of error with every rate; otherwise a change between two checks may be noise.
Should I trust a vendor's own comparison page?+
Only where every claim about a competitor is dated and sourced, and where the vendor says what it did not verify. This page is written by a vendor. Every competitor fact on it is a published price or plan from that vendor's own site, with the date we read it, and everything else is marked not verified.
How much should an AI citation tracker cost?+
Published entry prices in this category run from $29 to $99 a month, with top tiers from $149 to over $1,000. The price is mostly a function of how many prompts, engines and repeated runs you get, so compare the number of sampled answers per month, not the headline.
See which pages answer your category — and what they say about you.
Check your site →