AI Visibility Software
← Blog
Published

The AI Visibility Software Evaluation Scorecard: 40 RFP Questions With Scoring Weights (Free Template)

A neutral, weighted RFP scorecard for AI visibility software: 40 questions across eight categories, with point values and red-flag answers to watch for.

Bottom line

Score AI visibility software RFPs across eight weighted categories: engine coverage, prompt methodology, measurement rigor, action capability, reporting, security, support, and pricing. Run the 40 questions below, weight each category from five to 20 points, and disqualify any vendor whose answer trips a listed red flag.

Last updated October 2026. Pricing, plan structure, and vendor capabilities below reflect published vendor pages at the time of writing.

Most vendor scorecards get written by the vendor’s sales team, and it shows. The questions favor whatever that vendor already does well and skip the categories where a competitor would win. A 2026 G2 “Answer Economy” survey found that 51% of B2B software buyers now start purchase research in an AI chatbot, up from 29% a year earlier, which makes the tooling that measures this channel a real procurement decision, not a side purchase.

This scorecard is vendor-neutral. It asks the same 40 questions of every platform on your shortlist, weighs the answers the same way, and flags the specific responses that should worry a buyer regardless of which vendor gives them. Use it as-is, or adjust the category weights to match your team’s priorities before you send it out.

How to use this scorecard

Send the eight category tables below to every vendor on your shortlist, or work through a live demo with them open. For each question, award the full point value for a strong, evidence-backed answer, half credit for a vague or partial one, and zero for a red flag.

Total the points inside each category, then add the eight category totals for a score out of 100. A vendor that cannot answer a question at all scores zero on it, the same as a red flag.

At a glance: the 8 categories

CategoryWeightWhat it tests
Engine coverage and data freshness20Which engines it tracks, and how current the data is
Prompt methodology and sourcing15Where prompts come from and how transparent the evidence trail is
Metrics and measurement rigor15Whether the numbers hold up to scrutiny
Action and execution capability15Whether it ships fixes or just shows gaps
Reporting and integrations10Whether the data reaches the tools your team already uses
Security, compliance, and governance10Whether procurement and legal can sign off
Support and onboarding10How fast you get to a usable first result
Pricing and contract terms5Whether the true cost is visible before you sign

For a fuller side-by-side comparison of pricing and engine coverage beyond RFP questions, see the full tool ranking.

Category 1: Engine coverage and data freshness (20 points)

This category carries the most weight because a monitoring gap here undermines every other metric on the scorecard. A platform that tracks ChatGPT but skips Google AI Overviews is answering half your buyers’ question. Profound covers nine-plus engines at its Growth tier and prices depth into the product. Peec AI ships a lean three-model starting point and prices additional engines as add-ons. Ask every vendor to show both, not just describe them.

#Ask thisPoints
1Which AI engines does the platform track today, and is Google AI Overviews included by default or gated behind an add-on?4
2How often is data refreshed for each tracked engine, and does refresh speed change by plan tier?4
3Are Google AI Overviews and Gemini tracked as separate, distinctly labeled surfaces?4
4What happens to historical trend data if the vendor adds or drops an engine?4
5Can we add a new engine to an existing prompt set without repeating full onboarding?4

Category 2: Prompt methodology and sourcing (15 points)

Two vendors can report the exact same share-of-voice number from two completely different prompt sets, and only one of those sets reflects how your buyers actually search. Ahrefs Brand Radar builds its prompt corpus from a database of more than 405 million real search queries pulled across six AI platforms. Semrush’s AI Visibility Toolkit auto-discovers prompts from the keyword universe already inside your Semrush account. Neither approach is wrong, but they produce different prompt sets, and your RFP should force each vendor to name which one you are buying before you sign.

#Ask thisPoints
6Are prompts sourced from real search or chat query data, written manually by the buyer, or auto-discovered from an existing keyword list?3
7How many prompt variants per intent does the platform track by default, and is there a hard cap on the base plan?3
8Does the platform show the literal prompt text and the engine response behind a scored citation, or only an aggregated percentage?3
9How does the platform handle prompt drift, meaning it retires prompts that no longer reflect how buyers phrase the question?3
10Is competitor-level prompt data visible in the same interface, or only our own brand’s results?3

Category 3: Metrics and measurement rigor (15 points)

A single-run dashboard reports noise, not a signal, because answer engines are probabilistic and one response is one sample from a distribution. This category checks whether the vendor’s numbers would survive an internal audit before your team presents them upward.

#Ask thisPoints
11Does the platform report share of voice, brand mentions, citation rate, and sentiment as 4 distinct metrics, or collapse them into one score?3
12How many runs per prompt does the platform sample before reporting a percentage?3
13Can we export raw, prompt-level data, or only the aggregated dashboard view?3
14Does sentiment scoring distinguish positive, neutral, and negative, or return a binary flag?3
15How does the platform flag factual inaccuracies about our brand, separate from sentiment?3

Category 4: Action and execution capability (15 points)

Most platforms stop at the dashboard: they show you the gap and hand the fix to your team. Temso closes that loop from $89/mo, its dashboard turns a visibility gap into a ranked fix list and a content brief inside the same subscription, and its AI SEO agent can execute a subset of those recommendations directly, with a human review step before anything publishes. Ask every vendor whether “action” means a generic checklist or a fix tied to your actual gap.

#Ask thisPoints
16When the platform finds a visibility gap, does it produce a prioritized, ranked fix, or stop at the dashboard?3
17Can the platform generate a content brief or draft tied to a specific prompt where a competitor is winning?3
18Does the platform address citation gaps on third-party sources, or only owned-content recommendations?3
19Does the platform automate any part of execution, and exactly what happens without human review?3
20How does the platform confirm whether a shipped fix actually closed the gap over the following weeks?3

Category 5: Reporting and integrations (10 points)

Data that only lives inside a vendor’s dashboard rarely reaches the people who need to act on it. This category checks whether the platform pushes evidence into the tools your team already runs. Peec AI, for example, ships native connectors into Slack and BI tools rather than leaving export-and-import as homework for the analyst.

#Ask thisPoints
21Can reports be scheduled and exported as PDF, CSV, or via API without manual copy-paste?2
22Does the platform integrate with tools you already use, such as Slack, a BI tool, or GA4?2
23Can you build a separate dashboard view for an executive summary versus analyst-level detail?2
24Is a public API included in the base price, or gated behind a paid add-on?2
25Can multiple brands or properties be managed from one account without duplicate logins?2

Category 6: Security, compliance, and governance (10 points)

This is the category that stops a deal outright, not just loses points. Get every answer here in writing before the contract, not during onboarding.

#Ask thisPoints
26Is the vendor SOC 2 Type II certified, and can they provide the report on request?2
27Does the platform support SSO and role-based access control?2
28Where is customer data hosted, and is there a documented retention and deletion policy?2
29Does the contract include a data processing agreement covering GDPR requirements?2
30Who owns the prompt set and historical data after cancellation, and can we export everything first?2

Category 7: Support and onboarding (10 points)

A tool with perfect data is worthless if your team never gets past setup. This category is where budget-tier platforms most often lose points, since dedicated support usually gets reserved for higher plans.

#Ask thisPoints
31What does day-one onboarding look like: a guided setup, a dedicated rep, or a self-serve wizard with no human contact?2
32What is the average time from signup to a usable first dashboard?2
33Is customer support included at every tier, or reserved for higher-priced plans?2
34Does the vendor offer training or documentation for a team without a dedicated AEO analyst?2
35What is the vendor’s process for disputing a mention or sentiment score you believe is wrong?2

Category 8: Pricing and contract terms (5 points)

The true cost of an AI visibility platform rarely shows up on the pricing page. Profound’s Starter tier covers ChatGPT only at $99/mo; real coverage sits at Growth for $399/mo. Semrush’s AI Visibility Toolkit is a $99/mo add-on that requires a paid Semrush Pro plan starting at $139.95/mo before it activates. Peec AI’s Starter plan runs €85/mo, with additional engines priced as add-ons from €30 to €140/mo each. Temso holds three flat tiers, $89, $199, and $499 per month, with unlimited projects, users, and recommendations on every one of them, plus 15% off on an annual commitment.

#Ask thisPoints
36Is pricing published, or does every quote require a sales call?1
37Are engines, users, or prompt volume metered as paid add-ons, or included in the base price?1
38Is there a free trial or pilot period, and does it require a credit card?1
39What is the cancellation policy, and is there a minimum contract term?1
40Does an annual commitment carry a disclosed discount, or is it negotiated case by case?1

Reading your results

A score above 85 with zero red flags in security or measurement rigor is a strong candidate. A score between 65 and 84 is workable if you can negotiate around the specific gaps a vendor showed. Below 65 signals problems across multiple categories, not just one weak area.

Treat category 6 differently from the rest. One unresolved security red flag should disqualify a vendor regardless of its total score, because a data-governance gap does not average out against a strong prompt methodology.

The default weights above fit a general buyer. An agency running several client accounts should raise reporting and integrations. A team with procurement and legal in the room should raise security and governance before the RFP goes out, not after responses land. The same weighted-category logic drives the scoring behind our tool ranking, documented in full on the methodology page.

Your scoring sheet

Copy this table into a spreadsheet, add a column per vendor, and fill in the category totals as responses come in.

CategoryMax pointsVendor AVendor BVendor C
1. Engine coverage and data freshness20
2. Prompt methodology and sourcing15
3. Metrics and measurement rigor15
4. Action and execution capability15
5. Reporting and integrations10
6. Security, compliance, and governance10
7. Support and onboarding10
8. Pricing and contract terms5
Total100

Run this against your own shortlist

Print the eight tables above, send them to every vendor on your list, and score the responses the same way for each one. The scorecard only works if you apply it evenly, so resist the urge to skip a category for the vendor you already like.

If you want a fast baseline before the RFP even goes out, Temso covers the engine, prompt, and action categories in one $89/mo plan with a free trial and no credit card required, giving your committee a working answer to compare every other response against.

FAQ

What should an AI visibility software RFP cover?

A complete RFP for AI visibility software covers eight areas: which AI engines the platform tracks and how fresh that data is, how it sources and manages prompts, how rigorous its metrics are, whether it ships fixes or just a dashboard, its reporting and integrations, its security and data governance, its onboarding and support, and its pricing and contract terms. Weighting engine coverage and measurement rigor the heaviest reflects where vendors differ most.

How many questions should I ask AI visibility vendors?

Forty questions across eight categories, five questions per category, is enough to separate serious platforms from thin ones without turning procurement into a multi-week project. Fewer than 20 questions misses too many failure modes; more than 50 slows down the buying committee without adding much signal.

What counts as a red flag in an AI visibility RFP response?

A red flag is any answer that hides the real cost or capability gap until after signup: a vendor that cannot show the literal prompt and response behind a scored percentage, a refresh cadence slower than weekly with no faster tier, per-engine add-on pricing not disclosed until after the demo, or no SOC 2 report and no timeline to get one. One category-6 (security) red flag should stop the process outright.

How do I weight the scorecard categories for my team?

The default weights (20, 15, 15, 15, 10, 10, 10, and 5 points) fit a general buyer. An agency running multiple client accounts should raise reporting and integrations. An enterprise team with procurement and legal in the room should raise security and governance. Reweight before you send the RFP, not after the vendor responses come in.

What total score should disqualify a vendor?

Below 65 out of 100 signals real gaps across multiple categories, not just one weak area. Between 65 and 84 is workable if you can negotiate around the specific gaps. Above 85 with zero red flags in security or measurement rigor is a strong candidate. Any unresolved red flag in the security and governance category should disqualify a vendor regardless of its total score.

Do Ahrefs Brand Radar and Semrush source prompts the same way?

No. Ahrefs Brand Radar builds its prompt corpus from a database of more than 405 million real search queries pulled across six AI platforms. Semrush's AI Visibility Toolkit auto-discovers prompts from the keyword universe already inside a customer's Semrush account. Both are legitimate starting points, but they produce different prompt sets, and an RFP answer should name which one a buyer is getting.