Most teams starting on AI visibility do prompt research for about five minutes: they type their brand name into ChatGPT, watch it describe them accurately, and declare victory. That tells you nothing. Nobody who hasn’t heard of you is typing your name — they’re describing a problem and asking which tool, service, or answer solves it. The real work here is reconstructing those problem-shaped questions before the buyer ever mentions you, then deciding which of them are worth measuring and optimizing for. Done well, it becomes the seed list for everything downstream: what you track, what you write, and how you know whether AI search is sending you anyone.
Why Prompt Research Isn’t Just Keyword Research
The instinct is to export your keyword list, and it’s half right — your keywords are a starting seed, not the answer. Keywords are compressed. Someone searching Google types “crm for small business” because the ten blue links do the rest of the reasoning. A prompt carries the reasoning inside it: “I run a five-person agency, we track clients in spreadsheets and keep dropping follow-ups — what CRM should we actually use and why?” That single query bundles context, constraints, intent, and an implicit ask for justification. This is the discipline of recovering those fuller utterances, because those are what people genuinely ask assistants, and those are what the model is answering when it decides whether to mention you.
The second difference is that prompts are conversational and chained. A searcher runs one query; an assistant user runs a thread — an opening question, a follow-up that narrows it, a “which of those is cheapest,” a “does it integrate with X.” Your visibility can live or die on the second or third turn, not the first. Any research that only captures cold openers misses where buying decisions actually get made.
The Uncomfortable Truth: There’s No Prompt Volume
Keyword research rests on a number — search volume — that tells you how many people want a thing. Prompts have no equivalent, and pretending otherwise is the fastest way to waste a quarter. The assistants don’t publish query data. Any “prompt volume” figure you see is modeled, inferred, or invented, and phrasing varies so widely that two people with the identical need will word it a dozen different ways. So the whole prioritization model has to change. Instead of ranking prompts by a volume estimate, you rank them by three things you can actually reason about: commercial intent (does answering this put you in front of a buyer), realistic surfacing odds (could a model plausibly cite you here today), and representativeness (does this prompt stand in for a whole family of similar phrasings). Accept that you’re working with a representative sample, not a census, and the exercise becomes tractable.
Where the Real Prompts Come From
Good ai prompt research is sourced from places where humans already phrase their needs in full sentences. Mine these, in rough order of signal quality:
- Sales and support transcripts. The questions prospects ask on calls, and the ones that flood support, are near-verbatim prompts. “How is this different from [competitor]” is a comparison prompt someone is running in ChatGPT right now.
- Your existing keyword clusters, expanded. Take each high-intent keyword and rewrite it as the full question a person would voice — add the constraint, the “for,” the “vs,” the “is it worth it.” One keyword seeds five to ten realistic prompts.
- Reddit, forums, and community threads. This is where people ask messy, honest questions — and it’s also disproportionately what assistants cite, so the phrasing there doubles as a surfacing clue.
- “People also ask” and autocomplete. Directional, but a cheap source of the sub-questions attached to a topic.
- Competitor comparison and alternative queries. “Best [category],” “[competitor] alternatives,” “is [competitor] worth it” — high commercial intent and easy to check whether you surface.
The goal at this stage is coverage, not precision. Cast wide, then cut.
Build a Prompt Taxonomy, Not a Flat List
A pile of two hundred prompts is useless until you sort it by intent, because you’ll act on each type differently. A workable taxonomy for geo prompt research has four tiers that map to the funnel:
- Problem-exploratory — “why does my site get traffic but no leads.” The buyer doesn’t know the category yet. Winning here is about being cited as a source in a broad answer.
- Solution / category — “what tools help track AI search visibility.” They know the category, not the players. This is where you fight to be named.
- Comparison — “X vs Y for a small agency.” High intent, and often where the assistant’s framing decides the sale.
- Branded / verification — “is [your brand] any good.” Lowest volume, highest intent — someone is checking you out. You want the model’s answer here to be accurate and flattering-because-true.
Tagging every prompt this way turns the exercise from a wish list into a map. It also stops you from over-investing in branded prompts (comforting, but you’re already winning them) at the expense of the category and comparison prompts where new buyers actually find you.
Cluster by Phrasing, Because Models Are Sensitive to Wording
Here’s the mechanism most guides skip: assistants can return meaningfully different answers to prompts that mean the same thing. “Best AI SEO tool” and “top software for AI search optimization” are the same intent, yet the model may cite a different set of sources for each because it’s pattern-matching on the exact tokens and the content that happens to align with them. So a single prompt is a fragile unit to track. Cluster three to five phrasings around each intent and treat the cluster as your unit of measurement. If you surface for one wording and vanish for the others, that’s not noise — it’s a content gap telling you which phrasings your pages don’t yet map to.
Prioritize Without Volume: The Scoring Rule
With no volume to sort by, use a simple three-factor score per prompt cluster, rated low/medium/high:
- Commercial intent — how close is this to a buying decision?
- Surfacing feasibility — given your current authority and content, could you realistically be cited here in the next quarter, or is this a two-year climb?
- Strategic fit — does winning this prompt reinforce the entity you want to be known for, or is it a tangent?
Track the clusters that score high on at least two axes; park the rest. The trap is chasing prestige prompts (“best marketing software”) where you have no realistic surfacing odds while ignoring the narrower comparison prompts you could own this month. Prioritization by feasibility, not ambition, is what separates a working program from a dashboard nobody trusts.
Validate Prompts by Actually Running Them
A prompt only earns a place on your tracking list after you’ve run it and seen what the assistant returns. Run each candidate cluster across the surfaces that matter to your audience — ChatGPT, Perplexity, Google’s AI Overviews, Gemini, Copilot — and record three things: whether you appear at all, who does appear (your real competitive set in AI answers, which is often not your SEO competitive set), and what sources the model cites to justify its answer. That citation list is gold: it’s a literal to-do list of the pages and domains you’d need to influence, appear alongside, or displace. This is also where you discover which prompts are even winnable — if every answer cites three entrenched authority sites and a Reddit thread, you now know the shape of the battle before you spend on it.
Turn Prompt Research Into a Tracking Set
Because there’s no query surface for prompts, AI visibility is invisible unless you deliberately instrument it — and that’s exactly the measurement layer SEO Rocket is built for. Its AI-visibility tracking runs your chosen prompt set on a schedule across ChatGPT, Gemini, Google AI Overviews, and Perplexity, and records how often your brand appears and gets cited, so an otherwise unmeasurable surface becomes a trend line you can report on. The output of your prompt research — the prioritized, clustered, validated list — is precisely the input that layer needs. Start narrow: fifteen to thirty prompt clusters that cover your core buying journeys beat two hundred you’ll never look at. You can always widen the set once you trust the signal.
Feed the Same Research Back Into Content
This research isn’t only for measurement — it’s a content brief in disguise. Every prompt where you fail to surface is a page (or a passage on an existing page) that doesn’t yet answer that question well enough to be quoted. The comparison prompts tell you which “X vs Y” pages to build; the category prompts tell you which definitive overviews to write; the exploratory prompts tell you which top-of-funnel explainers earn citations. This is where the rest of the workflow earns its keep: seeding prompts from real keyword data (SEO Rocket runs keyword research on live Ahrefs data), finding the topics where rivals surface and you don’t through competitor gap analysis, and drafting the answer with a validation-gated writer that enforces genuine depth rather than thin filler. Research the prompt, measure the gap, write the answer, re-measure — that loop is the entire program.
How Many Prompts Should You Track?
Fewer than you think. The failure mode is a sprawling list that decays into noise because no one can act on two hundred rows. A focused set of fifteen to thirty clusters, chosen for commercial intent and feasibility and refreshed quarterly, tells you more than an exhaustive one you never revisit. Add prompts when a new product, market, or competitor makes a genuine buying journey visible; retire prompts that turned out to have no commercial pull. And report the trend, not the snapshot — assistant answers vary run to run, so one bad check means nothing and one good check means less. If you’re reporting AI visibility to clients, a client dashboard that shows appearance and citation share moving over weeks is far more honest than a single lucky screenshot.
Common Mistakes That Waste the Effort
Three recurring errors sink these programs. First, tracking only branded prompts — reassuring, but you’re measuring a race you already lead instead of the category prompts where new buyers actually meet you. Second, single-phrasing tracking — one wording per intent hides the phrasing sensitivity that’s the whole point, so you get false confidence or false alarm. Third, treating the list as static — buying language shifts fast in AI search, and a prompt set you set up in January and never revisited is measuring last quarter’s market. Avoid those three and the discipline pays off; fall into them and you’ve built a dashboard that looks busy and teaches nothing.
Frequently Asked Questions
Is prompt research the same as keyword research?
No, though your keywords are a useful seed. Keywords are compressed search inputs; prompts are fuller, conversational utterances that carry context, constraints, and an implicit ask for justification — and unlike keywords, they have no reliable volume data. Prompt research recovers those full questions and prioritizes them by intent and feasibility rather than by a search-volume number.
How do I find the prompts people actually ask AI assistants?
Mine the places people already phrase needs in full sentences: sales and support transcripts, community threads on Reddit and forums, “people also ask” boxes, competitor comparison queries, and your own keyword clusters rewritten as complete questions. Then validate by running each candidate through ChatGPT, Perplexity, and AI Overviews to see who actually surfaces.
How many prompts should I track for AI visibility?
Start with fifteen to thirty prompt clusters that cover your core buying journeys, chosen for commercial intent and realistic surfacing odds. A small, well-chosen, quarterly-refreshed set that people actually act on beats an exhaustive list nobody revisits. Widen it only once you trust the signal.