Keyword research and discovery is the process of finding the queries your audience actually types, then filtering them down to the ones you can realistically win and grouping them into pages. Discovery is the expansion half — going from three seed terms to two thousand candidates. Research is the contraction half — cutting those two thousand down to forty pages worth building.
Most people do the first half and stop. That produces a spreadsheet, not a plan. The value is entirely in the filtering and grouping that comes after.
Start with seeds, not with a tool
Before you open anything, write down 5–15 seed terms by hand. These come from three places: how your customers describe the problem in their own words, the terms your sales or support conversations keep repeating, and the two or three competitors you actually lose deals to.
Seeds should be short and unglamorous — two or three words, no modifiers. “rank tracking,” not “best affordable rank tracking software for agencies.” The modifiers are what the discovery step generates for you. Starting too specific narrows the funnel before it opens.
Add your own site as a seed source. Export the queries you already get impressions for in Search Console, sorted by impressions with low click-through. Those are queries where Google already considers you relevant and you are not converting the visibility — often the fastest wins available.
Discovery: expanding the candidate set
Run each seed through a keyword tool and collect everything. At this stage you want breadth, not judgment. Five channels are worth pulling from.
- Related and matching terms from a keyword database, with volume, difficulty, CPC, global volume, and SERP features attached. SEO Rocket returns up to 150 ideas per search and supports multi-seed queries, so you can stack several seeds into one working set.
- Competitor organic keywords — everything a rival domain ranks for. This is usually the richest single source because those queries are already proven to convert traffic in your market.
- Content gap across several competitors at once, showing queries where they rank and you do not, with per-rival position columns so you can see who is weak on each.
- Google’s own surfaces — autocomplete, People Also Ask, related searches at the foot of the page. Free, current, and straight from the source.
- Community language — forums, review sites, support tickets. This is where you find the phrasings that keyword databases underreport because they are new or low-volume.
Expect 500 to 3,000 raw candidates for a normal topic area. Do not clean as you go; collect first, then filter in one pass.
Filtering: the cuts that matter
Work through these in order. Each one removes a large chunk cheaply.
- Wrong country index. A site serving Singapore queried against the US index returns near-useless data. Set the market first — this single setting invalidates entire research sessions when it is wrong.
- Competitor brand terms. You will not rank for a rival’s name and you would not want the traffic anyway.
- Irrelevant intent. Free-tool seekers when you sell enterprise, job seekers, students, DIY searchers when you sell services.
- Volume floor. Cut below 10 searches a month unless the term is unusually commercial. Remember volumes are modeled roughly as 12-month averages, so 20 versus 30 is the same number.
- Unwinnable difficulty. Not by score alone — check the results page. A high difficulty score with a forum thread in position seven is more winnable than a moderate score where every result is a major brand.
What remains should be a few hundred keywords you would genuinely be happy to rank for. That is the research output.
Scoring what is left
Rank the survivors so you know what to build first. A simple, honest score beats a complicated one.
Weigh commercial intent heaviest — CPC is a decent proxy, since advertisers do not pay for worthless clicks. Then winnability, judged by the weakest page-one result rather than the strongest: how many referring domains point at it, how deep is it, could you beat it in a quarter. Then volume, which matters least of the three and misleads most often. A 90-search query where the searcher is ready to buy outperforms a 5,000-search query full of casual readers.
Factor in SERP features too. If a query has an AI Overview, four ads, and a shopping carousel above the first organic result, the traffic at position one is a fraction of what raw volume implies. A clean results page at the same volume is worth considerably more.
Grouping into pages
Keywords do not map one-to-one onto pages. Group by SERP overlap: pull the top ten URLs for each keyword and put keywords together when they share three or more of the same results. Google has already decided those queries want the same page, and fighting that decision is how you end up with four near-identical posts cannibalizing each other.
Most healthy groups hold 5–40 keywords. Above 50, look for two intents accidentally merged. Below 5, the term probably belongs inside a broader page. For each group, name a primary keyword — best volume-to-difficulty ratio, not highest volume — assign a page type based on what currently ranks, and list the specific questions the page must answer.
That output is a brief. It should be handed to a writer or a tool without further explanation.
One more thing belongs in every brief: the differentiator. Name the specific element your page will carry that nothing on page one currently has — original numbers, a worked example, a comparison table, a screenshot from real use, or a clear opinion. Pages that only recombine what already ranks tend to settle around position eleven and stay there, no matter how clean the research behind them was.
Keeping the research alive
Research decays. Save the surviving keywords into a project pool rather than a dead spreadsheet, so the same set feeds your content production and your rank tracking. Then revisit on a schedule: priority clusters quarterly, the long tail every six months, and anything immediately after a confirmed core update where you saw movement.
Watch Search Console continuously for terms you never researched. Pages accumulate rankings for queries nobody predicted, and those are free signals about what your audience actually wants. Feed them back in as new seeds.
Where the numbers are and are not reliable
Be clear-eyed about the data. Third-party volume, difficulty, and position figures are estimates built from models and periodic crawls, not readings from Google. Use them for relative comparison — this term is bigger than that one, this competitor is weaker — and never as precise truth. For your own site, Search Console impressions and clicks always win.
SEO Rocket keeps both in one place at a flat US$50 a month: a multi-seed keyword explorer with country-specific indexes, competitor content gap across up to five rivals, a project keyword pool that feeds the AI writer and the tracker, and top-100 rank tracking with movement deltas alongside connected Search Console and GA4. Filtering and sorting with CSV export is free after one query, so you can do the contraction half of the work wherever you prefer. Run one full seed-to-cluster pass on your main topic this week; the page count that comes out the other side is usually a lot smaller and a lot more useful than the list you started with.