Part 2 · Practice · Updated September 13, 2026 · 7 min read
Building a Prompt Set That Represents Real Buyers
Every GEO number you will ever look at is an answer to a question somebody chose to ask. Choose shallow questions and you get a shallow measurement, no matter how good the tool is. This chapter is about building a prompt set that represents your real buyers instead of your own guesses.
Why prompt choice is the whole measurement
Traditional SEO starts from keyword data: someone else already counted what people search. AI engines publish no such data. There is no query report for ChatGPT. So every GEO program starts by deciding which questions to track, and that decision quietly determines everything downstream.
Get it wrong in either direction and the numbers lie to you. Track only prompts that contain your brand name and you will conclude you are famous. Track only generic questions no buyer would ask and you will conclude AI search does not matter for you. The prompt set is the sample; the sample decides what the study can see.
Start from buying questions, not keywords
A prompt is not a keyword. People type “crm for startups pricing” into Google. They ask ChatGPT “we’re a 15-person startup, what CRM should we use and what will it cost us?” The intent is the same; the shape is completely different. Longer, conversational, loaded with context.
So do not import your SEO keyword list. Instead, write down the questions a buyer would actually ask at each stage:
- Awareness: “How do companies figure out whether AI tools mention their brand?”
- Comparison: “What are the best AI visibility tools for a mid-size marketing team?”
- Decision: “Is [category leader] worth the price, or is there a better alternative?”
Group these into topics: coherent clusters of buying intent, such as “AI visibility measurement”, “GEO for agencies”, “alternatives to [competitor]”. Topics are the unit you will report on later, so pick ones a leadership team would recognize.
One question, many phrasings
Here is the part most first attempts miss: buyers do not share a script. Ten people with the same need will phrase the question ten ways, and AI engines answer different phrasings differently. A brand can be recommended confidently for “best GEO tools” and be invisible for “how do I track my brand in ChatGPT”, even though both come from the same buyer.
So each topic needs several phrasings: distinct, natural wordings of the same underlying question. Not synonyms swapped mechanically, but the way a blunt person, a detailed person and a skeptical person would each ask it. If you only ever ask one phrasing, you are not measuring the topic. You are measuring one sentence.
Who is asking changes the answer
Engines tailor answers to context the asker provides: role, company size, industry, geography, language. A question asked as a startup founder gets a different shortlist than the same question asked as an enterprise procurement lead.
That context is a persona. Decide which two or three personas actually buy from you and phrase prompts the way each of them would, with the context they would naturally include. If your buyers ask in more than one language or market, that belongs here too. A brand that looks strong for English-speaking startup founders can be absent for the same question asked from another market.
Ask more than once, always
Ask an AI engine the same question twice and you will often get two different answers: different brands, different order, different citations. That is not a bug you can wait out. It is how these systems work.
The practical consequence is simple: a single reading of any prompt is a coin flip pretending to be a fact. If your brand shows up in three answers out of five, and a single check happened to catch answer four, you would write down “invisible” and be wrong. Repetitions, meaning running the same prompt several times in fresh sessions, are what separate a measurement from an anecdote. What you want to know is not “did we appear?” but “how often do we appear?”
Putting the set together
A workable prompt set is a small grid, not a long list:
| Dimension | What it captures | A sensible start |
|---|---|---|
| Topics | The buying questions that matter | 5 to 10 |
| Phrasings per topic | How differently real people ask | 2 to 4 |
| Personas | Who is asking, from where | 1 to 3 |
| Repetitions | How much answers move between runs | 2 or more |
| Engines | Where your buyers actually ask | 2 to 5 |
Multiply that out and even a modest set produces hundreds of answers per measurement cycle. That is the point. Coverage across phrasings, personas and repetitions is what makes the resulting numbers trustworthy; volume on a single phrasing is what makes them merely long.
The coverage question you cannot answer from inside
There is a harder problem sitting underneath all of this, and it deserves naming honestly. A prompt set built in a conference room measures the questions your team could think of. Your market asks questions your team has never heard, phrased in ways nobody at your company would phrase them, and it asks some of them a hundred times more often than others.
How much of that real question space does your set cover? Ten percent? Ninety? From inside the set, this is unanswerable. You cannot audit your own imagination with your own imagination. Answering it requires modeling the question space itself: generating and running prompt variations at industry scale and measuring how much of the territory any given selection actually spans. That is a tooling problem, not a workshop problem, and no amount of team diligence substitutes for it.
The practical posture: build the small honest set, act on what it shows, and hold every conclusion with the humility of unknown coverage. The set tells you how you perform on the questions in it. It cannot tell you the set was the right one.
Mistakes that quietly ruin the measurement
- Branded prompts only. Tracking just “what is [your brand]?” style questions measures how well you defend ground buyers already found, and nothing about discovery. Keep branded prompts as your defensive line, and pair them with unbranded ones where the engine chooses who to name.
- Keyword transplants. Pasting SEO keywords as prompts produces questions no human asks, and answers no buyer sees.
- One phrasing per topic. You will mistake sentence-level luck for topic-level strength, in both directions.
- One run per prompt. You will chase changes that are noise and miss changes that are real.
- A frozen list. Buying questions shift. Revisit the set quarterly; retire topics that stopped mattering and add ones that emerged.
Questions people ask about this
How many prompts do I need to start?
Start small: five topics with two or three phrasings each, run repeatedly, will teach you more than two hundred one-off prompts. But be clear about what a hand-built set can never give you. It measures the questions you thought of; whether those are the questions your market actually asks, in the proportions it asks them, is a different problem. Mapping an industry’s real question space takes generating and running prompts at volumes no team does by hand, then measuring how much of that space a given selection covers. A DIY set is the right way to start and the wrong thing to be certain from.
Should prompts include my brand name?
Yes, and not only for sentiment and accuracy checks. Branded prompts are your defensive line: when a buyer asks specifically about you, engines routinely name competitors in the same answer. Tracking branded prompts tells you who is being recommended inside your own territory, and whether that share is growing. Pair them with unbranded prompts, where the engine chooses who to mention at all. One set defends your ground; the other measures whether you win new ground.
Do I need different prompts per engine?
No. Keep the prompt set identical across engines so differences in the answers reflect the engines, not your wording.