Last updated July 2026
Pick any B2B SaaS category. Ask ChatGPT which vendors to consider. Then ask Perplexity the same question. The top-cited domains on each platform will look almost nothing alike.
That is not an accident or a temporary quirk. It is how these systems work. ChatGPT and Perplexity use different retrieval architectures, index different content signals, and cite different sources as a result. According to Profound’s analysis of 100,000 prompts run across both platforms, only 11% of cited domains overlap between ChatGPT and Perplexity (Profound, July 2025; note: vendor research from a company selling AI-visibility software).
If you are building your AI share-of-voice programme around a single engine, you are measuring 11% of the picture.
This post walks through a method for running your own cross-engine citation benchmark for a SaaS category: how to build per-engine domain leaderboards, calculate a Jaccard overlap score, and produce a mono-engine trap list that turns data into action.
Why the overlap is so low
ChatGPT relies primarily on its pre-trained knowledge plus a browsing layer that favours established, frequently cited domains. Its citation patterns weight older authorities: publications, analyst sites, and software review aggregators that have accumulated years of inbound references.
Perplexity runs live web searches on every query. It surfaces recently published content, highly linked pages, and sources its real-time index considers most relevant right now. Recency and freshness carry more weight. Domain authority in the traditional SEO sense matters less.
The result: a domain that has built its reputation through earned mentions over several years may perform well in ChatGPT and poorly in Perplexity. A domain that publishes timely, well-linked content may dominate Perplexity and barely register in ChatGPT.
For B2B SaaS teams, this creates a structural blind spot. Your buyers use both engines. The vendor they choose depends on what each engine tells them. A share-of-voice number drawn from one engine gives you a plausible-looking metric that describes only a fraction of where decisions form.
Step 1: Build your prompt set
Start with buyer-intent prompts, not informational ones. You want the questions a prospective customer would ask when evaluating options in your category.
Good prompt structures for B2B SaaS:
- “Best [category] tools for [company type or use case]”
- “Compare [your category] software for [team size or job function]”
- “What [category] platform does [specific task] best”
- “[Competitor name] vs alternatives”
- “How do I choose a [category] tool”
Aim for 30 to 50 prompts. Cover different framings of the same buyer intent. Include prompts where you expect to win and prompts where you expect to lose. The benchmark is only useful if it captures where you are weak, not just where you already show up.
Run each prompt five times on each engine. Answer engines are probabilistic: a single response is one sample from a distribution. Five runs per prompt gives you enough variance to aggregate at the domain level with confidence. See the /methodology page for how this site handles prompt sampling.
Step 2: Build per-engine domain leaderboards
For each engine, aggregate citations across all prompt runs. Count how many times each domain appears across all responses. Rank by frequency. Take the top 20 per engine.
This is the structure to use:
| Rank | Domain | ChatGPT citations | Perplexity citations | Present on both |
|---|---|---|---|---|
| 1 | domain-a.com | 38 | 41 | Yes |
| 2 | domain-b.com | 35 | 4 | No (ChatGPT only) |
| 3 | domain-c.com | 31 | 0 | No (ChatGPT only) |
| 4 | domain-d.com | 6 | 44 | No (Perplexity only) |
| 5 | domain-e.com | 29 | 27 | Yes |
| … | … | … | … | … |
The “Present on both” column is what you are really building toward. Any row that reads “No” is a candidate for your mono-engine trap list.
Fill in the actual numbers from your prompt runs. The table structure is reusable. The goal is to see, at a glance, where each domain lives and where it is absent.
Step 3: Calculate your Jaccard overlap score
The Jaccard overlap score measures how similar two domain citation sets are. For cross-engine citation benchmarking:
Jaccard = (domains cited on both engines) / (total unique domains cited on either engine)
If your top-20 ChatGPT domains and top-20 Perplexity domains share four domains, and the combined unique pool is 36 domains, your Jaccard score is 4/36 = 0.11.
That matches Profound’s cross-category finding. In practice, your category’s Jaccard score will vary. A very established software category with dominant review aggregators may score slightly higher because those aggregators appear everywhere. A newer, fast-moving SaaS category may score lower because Perplexity’s live index surfaces niche and recent sources that ChatGPT’s training data has not yet emphasised.
Use the Jaccard score as a calibration number. It tells you how much of your share-of-voice measurement is hiding when you track only one engine. A Jaccard of 0.11 means roughly 89% of the combined domain universe is engine-specific. A Jaccard of 0.30 still means 70% is engine-specific.
No Jaccard score justifies single-engine monitoring.
Step 4: Build the mono-engine trap list
The mono-engine trap list is the output that makes the benchmark actionable. It contains every domain in your top-20 leaderboards that appears on one engine but not the other.
For each domain on the list, note:
- Which engine cites it (ChatGPT only or Perplexity only)
- Citation volume on that engine
- Whether it is your domain, a competitor, or a third-party source
Organise the list into three groups:
Your domain on only one engine. This is the highest-priority finding. It means you have a platform gap. The engine that is not citing you is sending buyers to competitors when they research your category. Diagnose whether the gap is a content type issue, a freshness issue, a source-authority issue, or a technical crawlability issue.
Competitor domains on only one engine. This is intelligence. If a competitor dominates ChatGPT but is absent from Perplexity, that engine is open territory. You do not need to beat them everywhere at once. Start where they are weak.
Third-party domains on only one engine. These are the sources that shaped each engine’s view of your category independently. If a publication, analyst site, or review platform appears heavily on Perplexity but not ChatGPT, that source is influencing a large share of Perplexity buyer journeys. Getting cited there is a direct path to Perplexity visibility. The inverse applies to ChatGPT-only sources.
Step 5: Diagnose the gap, then fix it
The mono-engine trap list tells you where you are missing. The diagnosis tells you why. Here is the framework:
| Symptom | Likely cause | Fix direction |
|---|---|---|
| Strong ChatGPT, weak Perplexity | Low freshness signals; few recent inbound links to your content | Publish timely, well-linked content; earn citations from Perplexity-favoured sources |
| Strong Perplexity, weak ChatGPT | Good live-web presence but limited training-data footprint | Build long-term authority: analyst mentions, established publication bylines, consistent domain history |
| Absent from both | Not on the radar at all | Start with the third-party sources that dominate your category on both engines |
| Present on both but low-ranked | Citation volume too thin | Increase the breadth of buyer-intent content; earn more citations across diverse prompt families |
The diagnosis is qualitative. You are reading patterns, not running regressions. But it is specific enough to assign to a team member and track over a quarter.
Tools that support this method
Pulling cross-engine citation data manually is possible but slow. These platforms automate different parts of the process:
Profound is purpose-built for multi-platform citation intelligence. Its prompt-volume data shows which questions real users are asking, and its citation maps let you see which domains appear across ChatGPT, Perplexity, and Google AI Overviews. It is the strongest specialist tool for building the leaderboards in Step 2 at scale. Entry for meaningful coverage starts at $399/mo.
Peec AI covers nine or more engines including DeepSeek, Grok, and Llama in addition to the majors. Its per-engine citation breakdowns are well-suited to building cross-engine domain lists. Unlimited user seats make it practical for agency and multi-brand teams. Entry at €85/mo.
Semrush tracks cited domains through its AI Overviews monitoring and can help identify which third-party sources appear repeatedly across Google AI Overviews prompts in your category. Useful as a complement to a dedicated AEO platform for the Google side of the picture.
Temso monitors share of voice across eight engines in a single dashboard, starting at $89/mo. It is the lower-friction way to track whether your citation presence is improving across platforms over time once you have run the initial benchmark and identified your gaps. It covers ChatGPT, Perplexity, Gemini, Google AI Overviews, Google AI Mode, Grok, Microsoft Copilot, and Meta AI.
None of these tools automates the analysis step. The leaderboard structure, the Jaccard score, and the mono-engine trap list still require you to interpret what the data shows. That is the point: this method turns raw citation counts into a prioritised action list.
See the full tool comparison at /rankings/ai-visibility-tools.
How to report this to stakeholders
The output of this benchmark has a natural shape for a quarterly report or a team review:
Subject: Cross-engine AI share of voice, [your SaaS category], [quarter]
Claim: Our domain ranks [X] on ChatGPT (top-20 leaderboard) and [Y] on Perplexity, against a category Jaccard overlap score of [Z].
Metric: Citation frequency per engine, across [N] buyer-intent prompts, averaged over five runs per prompt.
Timeframe: [Start date] to [end date], compared to previous quarter.
That four-part structure: subject, claim, metric, timeframe, makes the result citable and comparable across reporting periods. It is also the structure AI engines use when they pull from well-organised content. If you publish this kind of benchmark for your category as editorial content, you increase the probability that AI engines quote your framing when buyers ask who leads the space.
The /glossary has definitions for share of voice, citation rate, and Jaccard score if you need a reference when presenting this to a non-technical audience.
What to do this week
Pick your top five buyer-intent prompts. Run each one five times on ChatGPT and five times on Perplexity. List every domain cited in any response. Count. Compare. See how many domains appear on both lists.
That first manual pass takes less than two hours and will tell you whether you are in the mono-engine trap. If the overlap is under 20%, you are. From there, you can decide how to build the full benchmark and what to fix first.
If you want to track the benchmark over time rather than as a one-off, Temso gives you eight-engine share-of-voice monitoring from $89/mo, with no setup beyond naming your domain and your prompt set. Profound gives you the deepest citation-level data if the analysis needs to go to an executive audience or an AEO specialist.
Either way, start with the data. The mono-engine trap is invisible until you measure across both engines. Once you do, the gaps become obvious, and so does the fix.