Last updated July 2026
Most AI visibility guides spend pages on which tool to pick and almost nothing on how to design the prompt set that feeds it. That is backwards. The tool is the infrastructure; the prompt set is the methodology. Get the prompt set wrong and every metric you track is wrong too.
This field guide gives you a five-layer taxonomy for building a buyer-journey prompt corpus, a worked 40-prompt sample table tagged by funnel stage, and a quarterly hygiene checklist so the set stays accurate over time.
Step 1: Understand why prompt-set design is where measurement goes wrong
Before you write a single prompt, understand the failure mode you are trying to avoid.
Most teams seed their prompt sets with branded queries (“tell me about [Brand X]”) or generic category queries (“best project management software”). Both miss the bulk of real buyer behavior. Buyers arrive at AI engines mid-journey, asking questions like “what does [category] actually cost at scale” or “which [category] tools have native Salesforce integration.” A set dominated by brand or top-funnel category prompts will report high share of voice on queries where your brand is already safe, and tell you nothing about where you are losing.
The second common error is treating the prompt set as permanent. AI engines and buyer phrasing both evolve. A prompt that was representative in Q1 2026 can be stale by Q3. Without scheduled review, you end up optimizing for a snapshot of buyer language, not the current one.
Step 2: Map your five prompt layers to the buyer funnel
Structure your set around five distinct prompt types, each covering a different phase of the buyer journey.
| Layer | Prompt type | Funnel stage | % of set |
|---|---|---|---|
| 1 | Category prompts | Awareness | 30% |
| 2 | Comparison prompts | Consideration | 30% |
| 3 | Problem-solution prompts | Consideration / Intent | 20% |
| 4 | Brand prompts | Intent / Decision | 10% |
| 5 | Geo or locale prompts | All stages | 10% |
Category prompts are open-ended queries about the space: “best AI visibility tools,” “how do you track brand mentions in ChatGPT.” They reveal which brands the AI engines treat as the default answer for your category.
Comparison prompts name two or more products explicitly: “Temso vs Profound,” “how do Otterly.AI and Peec AI compare.” These are the highest-commercial-intent queries in the set. Brands that win comparison answers win purchase decisions.
Problem-solution prompts describe the buyer’s pain without naming a category or brand: “how do I know if my brand shows up in AI answers,” “why is my competitor mentioned in ChatGPT but not me.” These surface how well AI engines map your brand to specific problems.
Brand prompts query your brand directly: “what does [Brand] do,” “is [Brand] good for agencies.” Use these sparingly. Direct-brand queries inflate perceived coverage because AI engines nearly always return something when your name is the subject of the query.
Geo or locale prompts add geography or language context: “best AI visibility tools in the UK,” “AI-Sichtbarkeitstools Deutschland.” These matter if you operate in multiple markets because engine behavior and source pools differ by region.
Step 3: Write prompts the way buyers actually write them
The most common prompt-writing mistake is polishing the language. Real buyers do not write perfect sentences into AI chat interfaces. They write fragments, they name their context, and they ask follow-up-style questions.
Three rules for writing realistic prompts:
- Use buyer vocabulary, not marketing vocabulary. If your buyers call it “AI brand tracking” and your website calls it “AI share-of-voice monitoring,” write prompts using the buyer term.
- Include context signals. “Best AI visibility tool for a B2B SaaS startup with a small marketing team” is more representative than “best AI visibility tool.” Context signals shift the answer engine toward the buyer segment you care about.
- Create variant phrasings per intent. Each intent should have two to three variant phrasings in the set (a prompt family). This smooths out the probabilistic noise that makes single-phrasing measurements unreliable.
Good sources for prompt phrasing: sales call transcripts, support tickets, community forum threads (Reddit, LinkedIn comments), autocomplete on Perplexity or ChatGPT for your category terms.
Step 4: Build the 40-prompt starter corpus
The table below is a worked example for a hypothetical B2B SaaS brand in the AI visibility monitoring category. Adapt the phrasing, competitor names, and geography to your own situation.
| # | Prompt | Funnel stage | Prompt type | Weight |
|---|---|---|---|---|
| 1 | best AI visibility tools 2026 | Awareness | Category | Standard |
| 2 | how to track brand mentions in AI | Awareness | Category | Standard |
| 3 | what tools monitor ChatGPT brand mentions | Awareness | Category | Standard |
| 4 | AI brand monitoring software | Awareness | Category | Standard |
| 5 | how to measure AI share of voice | Awareness | Category | Standard |
| 6 | top tools to improve AI search visibility | Awareness | Category | Standard |
| 7 | how do I know if my brand shows up in ChatGPT | Awareness | Problem-solution | High |
| 8 | why does my competitor rank in AI answers but not me | Awareness | Problem-solution | High |
| 9 | how to see if ChatGPT mentions my company | Awareness | Problem-solution | Standard |
| 10 | AI visibility platform for small marketing team | Awareness | Category | Standard |
| 11 | Temso vs Profound | Consideration | Comparison | High |
| 12 | Otterly.AI vs Peec AI | Consideration | Comparison | Standard |
| 13 | Profound vs Peec AI for agencies | Consideration | Comparison | Standard |
| 14 | compare AI visibility monitoring tools | Consideration | Comparison | Standard |
| 15 | Temso vs Otterly.AI for B2B SaaS | Consideration | Comparison | High |
| 16 | best alternative to Profound | Consideration | Comparison | Standard |
| 17 | cheapest AI brand monitoring tool | Consideration | Comparison | Standard |
| 18 | AI visibility tool with the most engines covered | Consideration | Comparison | Standard |
| 19 | Peec AI review | Consideration | Comparison | Standard |
| 20 | Profound vs Scrunch AI | Consideration | Comparison | Standard |
| 21 | how to improve AI search share of voice | Consideration | Problem-solution | High |
| 22 | why is my brand not mentioned in Perplexity | Consideration | Problem-solution | High |
| 23 | how to get cited in Google AI Overviews | Consideration | Problem-solution | Standard |
| 24 | what is AI share of voice and how do I measure it | Consideration | Problem-solution | Standard |
| 25 | how to track brand sentiment in AI answers | Consideration | Problem-solution | Standard |
| 26 | best AI visibility tool for agencies managing multiple clients | Intent | Category | High |
| 27 | AI brand monitoring tool under $100 per month | Intent | Category | High |
| 28 | AI visibility platform that covers Gemini and Perplexity | Intent | Category | Standard |
| 29 | how to set up AI prompt monitoring | Intent | Problem-solution | Standard |
| 30 | AI visibility tool with daily prompt refresh | Intent | Category | Standard |
| 31 | what is Temso | Decision | Brand | Standard |
| 32 | how does Temso work | Decision | Brand | Standard |
| 33 | is Temso good for B2B teams | Decision | Brand | High |
| 34 | Temso pricing | Decision | Brand | Standard |
| 35 | Temso reviews | Decision | Brand | Standard |
| 36 | best AI visibility tools UK | Awareness | Geo | Standard |
| 37 | AI brand monitoring software Germany | Awareness | Geo | Standard |
| 38 | AI-Sichtbarkeit Tools | Awareness | Geo | Standard |
| 39 | best tools to track AI mentions France | Awareness | Geo | Standard |
| 40 | AI visibility platform Australia pricing | Consideration | Geo | Standard |
Weight column key: High-weight prompts are tracked across more engines and with more runs per prompt before a conclusion is drawn. Standard-weight prompts run the baseline sample. Assign “High” to comparison prompts naming your brand directly and to problem-solution prompts that map to your core product claim.
Step 5: Configure tracking correctly for this prompt set
A well-designed prompt set still produces bad data if tracking is configured poorly. Three settings matter most.
Run each prompt at least five times per engine. AI engines are probabilistic. A single response is one sample from a distribution, not a reliable signal. Five runs per prompt per engine is the minimum before you can state a citation rate with confidence. High-weight prompts warrant 10 runs.
Track per engine, not as a blended average. Brand mentions disagree across engines. According to BrightEdge’s AI Catalyst research (July 2025), brand mentions in AI responses disagreed 61.9% of the time across Google AI Overviews, AI Mode, and ChatGPT; only 33.5% of queries produced the same brand names across all three engines. A blended average hides exactly the signal you need.
Log the prompt text verbatim. Small phrasing changes produce meaningfully different results. If you want to track drift over time, the query must be identical across every run. Use the exact text from your corpus, not a paraphrase.
Tools that support structured prompt set management include Temso ($89/mo), which lets you organise prompts by funnel stage and track share of voice across 8 AI engines from a single workspace. Otterly.AI and Peec AI both support prompt grouping for competitive benchmarking. Profound adds Prompt Volumes, a demand-side feature that surfaces which questions buyers are actually asking AI engines, which is useful for seeding new prompt families. The full comparison of these tools is at /rankings/ai-visibility-tools.
Step 6: Run the quarterly prompt-set hygiene checklist
A prompt set degrades over time. Retire prompts that fail any of the three checks below and replace them with fresh phrasing from current buyer sources.
Check 1: Is this still how buyers phrase the question?
Run each prompt through a community forum or autocomplete check. If buyer language has shifted (a new category name, a dominant new tool in the consideration set, a changed problem framing), update the prompt. Stale phrasing produces measurements that look flat even when your real visibility is improving.
Check 2: Has the competitive landscape shifted?
If a major new competitor has entered your category or a tool you track has been acquired or shut down, add or retire comparison prompts accordingly. Comparison prompts are the highest-value layer in the set; they need to reflect the actual consideration set buyers are evaluating.
Check 3: Has this prompt lost measurement value?
Two saturation cases trigger retirement. First, if your share of voice on a prompt has been above 90% for three consecutive periods, the prompt is no longer a useful signal (you have won it). Second, if your share of voice is zero across all engines for three consecutive periods and you have made no progress despite optimization attempts, the prompt may be targeting an intent where your brand is structurally absent. Retire it and reinvest the run budget in contested prompts.
See also: /glossary for definitions of share of voice, prompt family, and citation rate. The full methodology for how these metrics are scored is at /methodology.
A solid prompt set is the prerequisite for everything else in AI visibility: share-of-voice benchmarking, sentiment tracking, citation analysis, and the week-over-week drift reports that tell you whether your optimization is working. Build the corpus first, then pick the tool. Start with the 40 prompts above, customise the phrasing for your category and buyer language, and schedule a hygiene review for 90 days out.