LLM SEO tracking software measures how often your brand appears in answers from large language models — ChatGPT, Google AI Overviews, Gemini, Perplexity — by sending a fixed set of prompts on a schedule and recording what comes back. That is the entire mechanism. There is no ranking API behind it, no index to query, no position to look up.
Understanding that one fact changes how you evaluate every tool in the category and how much confidence you put on the charts they produce. Here is the honest version of how this works, what the numbers mean, and how to run it without wasting a quarter.
Why this is sampling, not ranking
Classic rank tracking works because a search results page is a stable, ordered list. Send a query from a given country, read positions one through one hundred, store the snapshot. Repeat tomorrow and the deltas are meaningful.
Language models produce nothing of the kind. An answer is generated fresh each time. Temperature, context, personalization, and model version all shift the output. Ask “best project management tool for agencies” three times and you may get three different brand lists in three different orders. So tracking software runs the prompt repeatedly and reports frequency: out of N runs across M prompts, your brand appeared X times. That is a sample statistic with real error bars, not a position.
The metrics that are actually reliable
Sampling still yields decision-grade data if you stick to signals that survive the noise.
- Mention rate — the share of tracked prompts where your brand name appears at all. The most stable metric in the category and the one to build reporting around.
- Citation rate — the share where one of your URLs is linked as a source. Lower than mention rate, and the one that sends actual referral traffic.
- Prompt-level detail — the exact questions that produced a mention. Far more useful than an aggregate score, because it tells you which topics you already own.
- Co-mentioned sources — which third-party domains keep appearing alongside you. That is your outreach target list, generated for free.
Metrics to distrust: any “AI rank” expressed as a whole number, any traffic figure attributed specifically to an AI surface, and any sentiment score derived from a handful of runs.
What no tracking software can see
Vendors rarely lead with the limitations, so here they are in one place. Google does not separate AI Overview or AI Mode impressions and clicks in Search Console — they are folded into ordinary Web search data, with no filter that splits them out. Clicks from AI surfaces largely arrive in GA4 looking like standard Google organic referrals. ChatGPT and Perplexity referrals are somewhat more visible in analytics, but they undercount badly, because most model answers never produce a click at all.
The practical implication: you cannot build a clean attribution model from AI visibility to revenue today. What you can build is a presence trend and a correlation story. Treat anyone selling AI-sourced pipeline dashboards with proportionate skepticism.
Building your prompt set
The prompt set is the whole experiment. Get it wrong and every downstream number is decoration.
- Write 20–40 prompts that sound like a person talking, not a keyword list. “What’s a good rank tracker for a small agency?” beats “rank tracker software”.
- Cover four intents: category discovery (“best X for Y”), comparison (“X vs Y”), evaluation (“is X worth the money”), and problem-first (“how do I stop my rankings bouncing around”).
- Include your brand name in a few prompts to test what models say about you when asked directly — that surfaces stale or wrong claims fast.
- Freeze the set. Adding prompts mid-quarter resets your baseline. Version it and change it deliberately, at quarter boundaries.
- Log a control competitor. Even a manual check of one rival gives your own trend line context.
Cadence, sample size, and reading the trend
Weekly is enough for most sites; daily adds cost and noise without adding insight. Aim for at least three runs per prompt per check so a single odd generation cannot swing the number. With 30 prompts and three runs, that is 90 samples per model per check — enough that a move from 15% to 25% mention rate is worth discussing, and a move from 15% to 17% is not.
Read quarters, not weeks. AI answer variance is wider than the ±2–3 position drift you already accept in classic rank tracking, and model updates can shift everything overnight for reasons that have nothing to do with your site. Annotate your chart when a major model version ships; otherwise you will spend a month explaining a change you did not cause.
What SEO Rocket reports, and what it does not
SEO Rocket includes AI visibility in the flat $50 per month workspace. It reports brand mention counts across ChatGPT, Google AI Overviews, Gemini, and Perplexity, and shows the real example questions behind each mention, with no setup required. That covers presence, coverage, and the prompt-level detail that tells you which topics you already win.
It does not report AI Mode positions — those do not exist — and it cannot isolate AI traffic inside Search Console, because Google does not expose it. Competitor share-of-voice for AI visibility is on the roadmap and is not shipped. If a head-to-head share chart against named rivals is a hard requirement this quarter, a dedicated AI monitoring platform is the right purchase; those typically start in the low hundreds per month.
What to look for when buying
Feature lists in this category look interchangeable. Four questions separate the tools quickly.
- How many models, and are they named? Coverage across ChatGPT, Google AI Overviews, Gemini, and Perplexity spans most of the surfaces your buyers use. Vague “20+ AI engines” claims usually mean thin sampling on each.
- Does it show the raw answers? A tool that reports a score without the underlying response is asking you to trust a black box. Example questions and quoted text are the difference between a metric and an insight.
- How many runs per prompt per check? One run is a coin flip. Three or more is a measurement.
- What does setup cost you? Some platforms require weeks of prompt configuration before the first data point. Others report defaults immediately and let you refine later.
Price accordingly. Dedicated monitoring platforms commonly start in the low hundreds per month; AI modules bolted onto traditional suites usually sit behind a higher tier of a plan that already runs roughly $100–150. Bundled coverage inside a workspace you pay for anyway is the cheapest way to get a trend line started.
Turning tracking into changes worth making
Data only earns its cost if it changes what you publish. A few patterns show up consistently in pages that get cited:
- Answer first. A direct, self-contained answer in the opening 100 words is far easier for a model to lift than a claim buried under three paragraphs of setup.
- Be specific. Numbers, dates, prices, and limits get quoted. Adjectives do not.
- Map headings to sub-questions. Retrieval works at passage level, so a heading that matches the question is worth more than a clever one.
- Earn third-party presence. When your tracking shows the same review roundups and forum threads cited for your category, get accurate information onto those pages. That is often faster than moving your own site.
- Fix crawlability. Pages that are slow, orphaned, or blocked are poor retrieval candidates. A full-site crawl that shows the actual titles, URLs, and H1 text behind each issue finds those in an afternoon.
Run a fixed prompt set for two quarters, publish against the gaps it exposes, and keep your own Search Console data as the ground truth for total organic performance. If you would rather have that tracking sitting next to keyword research, rank tracking, and audits than paid for separately, SEO Rocket bundles it — just hold every number in the category, ours included, as a sample rather than a certainty.