Prompt set built from real questions
A researched, defensible set of prompts drawn from actual buyer language rather than a keyword list, kept small enough to run consistently.
AEO / AI citation tracking
There is no position one on a conversation. We track citation share against a defined prompt set instead, and we say plainly where that measurement is still immature.
Search engine optimization has thirty years of measurement infrastructure built under it. Rank trackers, click-through curves, impression data, all standardized and mostly reliable. Answer engine optimization has none of that yet. There is no stable position to track, because an assistant does not return a ranked list. It returns a synthesized paragraph, assembled fresh for that conversation, that may or may not name a source.
Citation tracking is the practice of building a defined, repeatable prompt set, running it against the major assistants on a schedule, and recording whether your brand appears, how it is described, and who else appears alongside it. It is the closest equivalent to rank tracking that currently exists for this surface, and it is a meaningfully weaker instrument than what SEO teams are used to. We think that should be said out loud instead of glossed over in a sales deck.
The core difficulty is non-determinism. Ask the same assistant the same question twice and you can get two different answers, with different sources cited or none at all. Model versions update without notice. A citation that appeared last week can disappear this week for no reason connected to anything on your site. Any measurement approach that treats a single run as ground truth is measuring noise.
What we do instead is treat citation tracking as a sampling problem. Multiple runs per prompt, tracked over time, looked at as a trend rather than a single data point. It gives directionally useful information about whether your visibility is improving. It does not give you a precise, defensible number the way a keyword rank does, and we will not pretend otherwise to make a report look tidier.
What it is
Five components, each addressing a different part of an honestly incomplete picture.
Prompt set design comes first and matters more than any tool used afterward. A prompt set built from a keyword list produces misleading results, because people phrase questions to an assistant differently than they phrase a search query. We build the set from real buyer questions, informed by sales conversations and support tickets, and we keep it small enough to run consistently rather than padding it with variations that add noise instead of signal.
Citation share is the headline metric: across the prompt set, in what proportion of runs is your brand named as a source, and in what position within the answer. We track this per assistant, since ChatGPT, Perplexity, Gemini, and AI Overviews draw on different retrieval systems and cite differently even for the same prompt.
Sentiment and framing matter beyond a binary citation. An assistant can name your brand accurately, name it inaccurately, or name it accurately but frame it as a secondary option behind a competitor. We record how you are described, not just whether you are mentioned, because a mischaracterization is worth catching even when it counts as a citation.
Competitor comparison and response volatility close the loop. We track which competitors appear on the same prompts and how consistently, and we flag prompts where the answer changes materially between runs, since those are the ones least safe to draw conclusions from. Referral traffic attribution sits alongside all of this as its own hard problem, covered separately below.
Fit
We would rather say no early than sell a program that cannot work.
Deliverables
A defined prompt set, run on a schedule, reported with its limitations attached.
A researched, defensible set of prompts drawn from actual buyer language rather than a keyword list, kept small enough to run consistently.
The proportion of runs in which you are cited, broken out by ChatGPT, Perplexity, Gemini, and AI Overviews, since each behaves differently.
How you are described when cited, not just whether you appear, including any inaccuracies worth correcting at the source.
Which named competitors appear on the same prompts and how their citation frequency compares to yours over the tracked period.
Prompts where the answer changes materially between repeated runs, marked as low-confidence rather than folded into the headline number.
A short document explaining sample size, run frequency, and exactly what the numbers can and cannot support, attached to every report.
How we run it
Four stages, repeated on a monthly cadence rather than run once.
We build the initial set from sales calls, support tickets, and category research, then review it with you before the first baseline run.
The full prompt set runs multiple times per assistant to establish a starting picture and flag which prompts are volatile from the outset.
The same set runs on a fixed monthly schedule, with results compared against baseline rather than treated as isolated snapshots.
We separate stable trends from noisy single-run results and report citation share, sentiment, and competitor movement with methodology attached.
Prompts where a competitor is consistently cited and you are not become inputs to the content and schema work under other AEO programs.
We build a defined prompt set from real buyer questions, run it repeatedly across assistants on a schedule, and track citation share, sentiment, and competitor comparison as a trend over time rather than a single precise position.
No, and we say so directly. There is no industry-standard tool with the reliability of a search rank tracker, model behavior changes without notice, and referral attribution from assistant traffic is still genuinely difficult. We report trends with methodology attached rather than presenting a single number as more certain than it is.
That immaturity is a reason to measure carefully, not a reason to skip measurement, since directional data is still more useful than none.
Where this connects
It closes the loop on structure and access work happening elsewhere in the pillar.
Citation tracking is diagnostic, not generative, so it depends on the answer-first pages built under AI search visibility actually existing to be cited.
A full audit of access, schema, and entity signals should happen before tracking starts, which is the scope of the answer engine audit and its baseline findings.
Which entity actually gets named in a citation depends on the disambiguation work done in entity optimization since an unresolved entity is harder for a model to cite consistently.
See how tracking fits alongside the other five programs on the AEO hub where the full sequence is explained.
Referral traffic attribution from assistant surfaces is its own measurement problem, handled under analytics so assistant-driven visits are not miscounted as direct traffic.
Questions
We run a baseline prompt set before proposing any ongoing tracking program.