Direct answer: use AI keyword research for expansion, normalization, classification, and cluster hypotheses. Do not use it as the source of search volume, ranking difficulty, current search intent, or business value. The reliable output is a short, evidence-backed queue of query families. It is not a list of hundreds of phrases ready to become pages.
This workflow stops at research prioritization. The topic cluster planning guide makes the later URL and internal-link decisions. If two live pages already appear for the same query family, use the Search Console cannibalization workflow instead of treating the problem as fresh keyword research.
What AI can and cannot do in keyword research
| Job | Good use of AI | Evidence still required |
|---|---|---|
| Expand language | Suggest wording, modifiers, questions, and adjacent jobs from supplied seeds | Customer language, Search Console queries, sales notes, and support records |
| Clean a list | Normalize spelling, flag duplicates, and preserve a source reference | A human check that meaningful distinctions were not erased |
| Propose clusters | Group phrases by the reader's task and explain each proposed boundary | Current result-page review and a collision test before assigning URLs |
| Classify intent | Suggest informational, comparison, transactional, local, or navigational labels | Manual review because one phrase can support several situations |
| Estimate demand | Prepare candidates for tool checks | Search Console, Keyword Planner, Trends, and real customer frequency |
| Prioritize | Apply a scoring rule consistently to evidence you provide | Business fit, service capacity, proof, and the next useful action |
Never ask the model to fill missing metrics. A plausible search-volume, CPC, trend, or difficulty number is still fabricated if it did not come from a named data source. Keep missing values blank and record what must be checked.
Step 1: begin with customer language, not an AI blank page
Collect the words people already use when a real problem is present. Useful sources include Search Console queries, sales-call summaries, contact forms, support tickets, internal site search, reviews, proposals, and objections recorded by staff. A small list of genuine questions is a stronger seed than a generic request for every keyword in an industry.
Remove names, email addresses, phone numbers, account details, confidential pricing, and customer-specific context before sending notes to an external AI service. Keep a neutral source label such as sales-call-07 or support-theme-refunds so every suggestion can be traced without exposing a person.
Start a research sheet with these fields:
- Raw phrase: the wording as received, including awkward language.
- Source: Search Console, customer call, ticket, review, or staff observation.
- Offering and audience: the service or product and the person making the decision.
- Problem stage: learning, comparing, buying, implementing, or troubleshooting.
- Location or constraint: market, platform, budget, company size, or deadline when relevant.
- Available proof: experience, screenshot, example, primary source, or measured result.
Step 2: expand candidates with a constrained prompt
Give the model the seed table and ask for controlled variations. Require a reason for every row and forbid invented metrics. The prompt below is deliberately narrower than “find keywords for my business.”
You are expanding a keyword research seed list.
Use only the business context and raw customer phrases below.
For each useful candidate, return:
- candidate_phrase
- source_seed
- reader_job
- likely_stage
- modifier_type
- ambiguity_or_missing_context
- evidence_needed
Rules:
1. Preserve meaningful customer wording.
2. Do not invent search volume, CPC, trend, competition, or rankings.
3. Do not turn simple synonyms or singular/plural variants into separate ideas.
4. Include a candidate only when it represents a real question, decision, or task.
5. Mark weak extrapolations as hypothesis.
Business context:
[paste a short factual description]
Sanitized seed rows:
[paste the seed table]
Review the result before continuing. Remove phrases that do not describe a useful reader job, promises the business cannot support, and variations created only by swapping adjectives. Keep the original source beside every surviving row.
Step 3: normalize and cluster by task
Clean obvious duplicates without destroying intent. Lowercase and trim whitespace for comparison, standardize spelling, and flag singular or plural variants. Preserve the original phrase in a separate column. Location, audience, platform, and problem-stage modifiers can change the job, so do not strip them automatically.
Then ask AI to propose clusters. A cluster is a shared reader task, not a bag of phrases containing the same noun.
Group these candidates by the reader's primary job.
For every proposed cluster, return:
- cluster_name
- one-sentence reader_job
- included candidate IDs
- phrases that look similar but must stay separate
- boundary_reason
- confidence: high, medium, or low
Do not merge candidates when they differ by:
- audience or decision maker
- problem stage
- local versus non-local need
- comparison versus implementation task
- troubleshooting versus general education
If the evidence is insufficient, place the candidate in "manual review".
Do not decide how many URLs to publish.
The final instruction matters. Query grouping happens before page architecture. One cluster may become a section, an existing-page update, a new guide, a service-page addition, or nothing at all.
Step 4: validate every cluster with named evidence
| Source | Best question it answers | Limit to record |
|---|---|---|
| Customer and sales records | Do real prospects use this language or need this decision? | Small samples reflect the customers you already reach |
| Google Search Console | Which queries already produce impressions or clicks for this site and page? | Anonymous queries are omitted and the table exposes only top rows |
| Google Ads Keyword Planner | Is there approximate demand in a selected location and period? | Volumes are rounded, close variants are grouped, and advertiser competition is not SEO difficulty |
| Google Trends | Is relative interest rising, falling, seasonal, or regionally different? | Values are normalized from 0 to 100 and do not show absolute search volume |
| Current search results | What jobs and result types does the wording currently surface? | Results vary by location, time, language, and personalization context |
| Business evidence | Can this site give a useful, specific, supportable answer? | A keyword can have demand and still be a poor fit for the business |
Search Console: first-party demand already touching your site
Start with the Performance report when the site already has data. Filter a relevant page, inspect its queries, and use a regular expression only when a well-defined phrase family needs grouping. Google omits anonymized queries for privacy and limits the table to top rows, so a visible query list is not the complete universe of searches. Bulk export provides the most complete query dataset when the property has enough data and the workflow justifies it.
Existing impressions are evidence of eligibility, not proof that a new URL is needed. Record the current page beside the query family. That single field helps prevent a research idea from competing with a page that already performs the job.
Keyword Planner: approximate market demand, not an organic forecast
Set the correct location, language, network, and date range. Google's average monthly searches include close variants and are rounded. Forecasts also depend on bid, budget, seasonality, and ad quality. Use the output as directional market evidence. Do not copy advertiser competition into a column called SEO difficulty, and do not present an ad forecast as expected organic traffic.
Google Trends: direction and seasonality, not volume
Compare candidates over the same geography and period. A Trends value of 100 marks the peak relative interest within that comparison, not one hundred searches. Low-volume phrases can display as zero. Compare a literal search term when wording matters and a topic when you need a broader concept across related terms or languages.
Manual result and evidence check
Open the live results for representative phrases and record what kind of task dominates: definition, step-by-step help, product category, comparison, local provider, tool, or troubleshooting. Do not imitate the top pages mechanically. The purpose is to test the cluster boundary and find what evidence a useful answer would need.
Step 5: score business fit before volume
Use a small scoring model that a second person can reproduce. The following ten-point model works well for a small content queue:
- Business fit, 0 to 3: no meaningful connection, adjacent, relevant, or directly connected to a service and qualified next step.
- Evidence, 0 to 3: unsupported, thin, adequate, or strong first-hand and primary-source support.
- Task clarity, 0 to 2: ambiguous, partly defined, or one clear reader job distinct from existing pages.
- Actionability, 0 to 2: passive information, a useful decision, or a practical next action the reader can complete.
| Query family | Business fit | Evidence | Task clarity | Actionability | Total | Decision |
|---|---|---|---|---|---|---|
| AI keyword research for a small service business | 3 | 3 | 2 | 2 | 10 | Prioritize |
| Free keyword list generator | 1 | 1 | 1 | 1 | 4 | Hold or answer briefly |
| Best SEO tool | 2 | 2 | 0 | 1 | 5 | Narrow the task first |
| Guaranteed number-one ranking keywords | 0 | 0 | 1 | 0 | 1 | Reject the premise |
Search demand can break a tie after the usefulness gates pass. It should not rescue a weak business fit, missing evidence, or a duplicate reader job. If a paid research platform would save time, the hands-on SE Ranking review shows what its keyword and competitor research can and cannot replace. A tool subscription does not remove the need for this evidence model.
Step 6: let AI apply the rule, not make the decision
Give the model only verified evidence and the scoring rubric. Require it to cite input fields and leave unsupported scores blank.
Score the validated query families using this rubric:
- business_fit: 0-3
- evidence_strength: 0-3
- task_clarity: 0-2
- actionability: 0-2
Use only the supplied evidence columns.
For each score, cite the exact input field that supports it.
If a field is missing, return "needs evidence" instead of guessing.
Flag any family already served by an existing URL.
Return a prioritized queue, a hold list, and a rejection list.
Do not recommend new URLs or page titles.
Review the ranking manually. A business constraint can override the total. For example, a high-scoring service topic should wait if the company cannot deliver that service, prove its claims, or handle the expected lead type.
The research handoff
The completed research sheet should contain one row per query family, not one row per wording variation:
- representative query family and included variations;
- reader job, audience, stage, and relevant location;
- original source and sanitized customer language;
- Search Console, Planner, Trends, and manual result evidence;
- business fit, available proof, score, and uncertainty;
- existing URL that may already serve the job;
- status: prioritize, hold, reject, or investigate.
Stop here. The next step is the content map, where the prioritized families are tested against existing URLs and given distinct page jobs. This separation reduces cannibalization because research does not silently create one page per keyword. If one family reveals a repeatable set of real entities rather than ordinary editorial topics, use the programmatic SEO quality-gate workflow to test the dataset before generating pages.
Common failure modes
- Starting from a generic AI prompt. The output reflects common web language, not evidence from your buyers.
- Treating every variation as a separate opportunity. Singular, plural, and close wording often belong to one task.
- Clustering by shared words. Similar vocabulary can hide different audiences, stages, and decisions.
- Inventing metrics. AI cannot know your current Search Console data or live tool settings unless you provide them.
- Sorting only by volume. A broad phrase can attract the wrong audience and consume evidence the business does not have.
- Ignoring the existing URL column. A promising phrase may belong in an update, not a new article.
- Uploading raw customer records. Sanitize notes and keep only the context needed for classification.
- Publishing the whole queue. A prioritized hypothesis still needs a page contract, sources, and a collision check.
Decision checklist
- Every candidate traces back to customer, search, market, or business evidence.
- The model is forbidden from inventing volume, competition, trend, or rankings.
- Clusters represent reader jobs, not repeated words.
- Search Console query limits and anonymized data are acknowledged.
- Keyword Planner settings and close-variant grouping are recorded.
- Google Trends is used for relative interest, not absolute demand.
- Business fit and available proof are scored before search volume.
- Every family records an existing URL that might already serve it.
- The output stops at prioritize, hold, reject, or investigate.
- URL architecture and live cannibalization remain separate workflows.
Sources checked
Tool behavior and reporting limits were verified from official Google documentation on August 2, 2026.
- Google Search Console: Performance report dimensions, anonymous queries, and table limits
- Google Search Console: common query and page analysis tasks
- Google Search Console: regular expression filters
- Google Ads Keyword Planner: historical metrics and forecast limits
- Google Trends: normalization, sampling, and interpretation
- Google Trends: search terms compared with topics
Need a defensible content queue?
I can turn customer questions, Search Console data, and current pages into a prioritized research map with clear evidence and overlap boundaries.
Request an SEO audit