Last updated September 2026
A brand disappears from a ChatGPT answer on a Tuesday morning. Someone notices, drops a screenshot in Slack, and the channel treats it like a five-alarm fire. Half the time, that reaction is correct. The other half, it is one probabilistic response among many, and the brand shows up fine on the next run.
Most teams have no way to tell which situation they are in. They have a monitoring dashboard, a gut feeling, and no defined line between “this is urgent” and “this can wait.” That gap is expensive. Real incidents sit unaddressed because they look like noise, and noise gets escalated because nobody set a bar in advance.
Site reliability teams solved this exact problem decades ago. A SEV-1 outage pages an engineer at 3 a.m. A SEV-4 bug goes in next sprint’s backlog. Nobody argues about which is which, because the criteria are written down before the pager ever goes off. AI visibility needs the same discipline. This ladder gives you four severity levels, each with a detection signal, a response window, an owner, and an action, so “our visibility dropped” stops being a feeling and starts being a classification.
The stakes are real. More than half of B2B software buyers, 51%, now begin their research inside an AI chatbot rather than a traditional search engine (G2, “The Answer Economy,” April 2026). A SEV-1 sitting open for a week is not a dashboard problem. It is buyers seeing a competitor on the exact query where you should have shown up first.
Why cross-engine confirmation matters before you panic
AI engines do not agree with each other nearly as often as most teams assume. BrightEdge’s AI Catalyst research found that brand mentions disagreed 61.9% of the time across Google AI Overviews, AI Mode, and ChatGPT, the three engines it tracked; only 33.5% of queries returned the same brand names across all three (BrightEdge AI Catalyst research, July 2025).
That fact should change how you triage. A brand missing from one engine, on one run, is close to the normal state of the system. A brand missing from the same money prompt on two or more engines, confirmed across repeated runs, is a different thing entirely. It is the disagreement rate collapsing into agreement, and agreement in this data is rare enough to matter.
This is the logic behind every threshold on the ladder below. The higher the severity, the more engines and the more repeated runs it takes to confirm the drop is real.
The ladder at a glance
Four rungs, ranked by how much revenue-relevant visibility is actually at risk.
SEV-1 CRITICAL Page someone now Money prompt gone, 2+ engines
SEV-2 HIGH Respond same day Money prompt gone, 1 engine; or sentiment turns negative, 2+ engines
SEV-3 MODERATE Fix this week Share of voice erodes on a tracked cluster; or one fact is wrong
SEV-4 LOW Log and monitor Sentiment drifts on one engine, secondary prompt only
A money prompt is a high-intent query, the kind closest to a buying decision: “best [category] tool,” “[brand] pricing,” “[brand A] vs. [brand B].” A secondary prompt is everything else: awareness-stage questions, feature explainers, general category education. The same drop matters far more on a money prompt than on a secondary one, which is why prompt type sets the ceiling on severity before engine count even comes into play.
How to define your money prompts before you need this ladder
The ladder only works if your money-prompt list exists before an incident hits. Building it takes an afternoon, not a quarter.
- Pull the last 10 to 20 questions your sales team says buyers actually asked during a demo or a win-loss call. Use their words, not your marketing copy.
- Add the two or three comparison prompts buyers run most: “[your brand] vs. [top competitor]” and “best [category] for [your core segment].”
- Cap the list at 10 to 15 prompts. A list you cannot check daily is a list nobody checks.
- Review the list every quarter. Buyer language shifts, and a stale prompt list will miss a real SEV-1 because you stopped tracking the query that mattered.
Everything else on this page, the tables, the runbook, the tool fit, assumes this list already exists. Skip it and every severity call downstream becomes a guess.
SEV-1: Critical
Your brand is gone from a money prompt on two or more of ChatGPT, Perplexity, Google AI Overviews, Gemini, or Microsoft Copilot, and the drop holds across repeated runs. This is the only level where you page someone the same hour you see it.
| Signal | Response window | Owner | Action |
|---|---|---|---|
| Brand absent from a money prompt on 2+ engines, confirmed across 3 runs per engine | Acknowledge in 1 hour; root cause in 4 hours | Growth or marketing lead, plus AEO analyst | Rerun the prompt, rule out a site outage or crawl block, check for a competitor’s new citation-worthy asset, escalate same day |
A SEV-1 that sits unaddressed for a week is not an anomaly anymore. It is your new baseline, and every day it stays open is a day buyers see a competitor instead of you on a query you should own.
SEV-2: High
Either your brand is gone from a money prompt on one engine only, or sentiment turns negative on two or more engines for a prompt that still names you. Both need a response inside the day, but neither needs a 3 a.m. page.
| Signal | Response window | Owner | Action |
|---|---|---|---|
| Money prompt gone on 1 engine; or sentiment flips negative on 2+ engines, confirmed across 3 runs | Acknowledge in 24 hours; resolve in 3 business days | AEO analyst or content lead | Diagnose the single-engine cause, draft a correction brief, check daily until it clears |
A single-engine dropout at SEV-2 is often the leading edge of a SEV-1. Treat it as a warning, not a coincidence, and check the other engines again before you close it.
SEV-3: Moderate
Share of voice on a tracked prompt cluster has been declining for two or more consecutive weeks, or one factual inaccuracy (an old price, a discontinued feature) has appeared on a single engine with no urgent revenue exposure attached to it.
| Signal | Response window | Owner | Action |
|---|---|---|---|
| Sustained week-over-week share-of-voice decline on a prompt cluster; or a factual error on 1 engine | Log within the week; fix in the next content sprint | Content or SEO team | Add to the backlog, correct the source content, request a recrawl, note it in the weekly review |
Most AI visibility work happens at this level. It is not urgent, but it compounds if nobody works the backlog.
SEV-4: Low
Sentiment drifts slightly on a single engine, on a prompt that is not a money prompt, or a competitor gains modest ground on a low-intent query. Nothing here needs a response today.
| Signal | Response window | Owner | Action |
|---|---|---|---|
| Minor sentiment drift, 1 engine, secondary prompt only | Log and monitor; revisit at the next monthly review | Whoever owns the monitoring dashboard | Note it in the tracker; act only if it persists past a month |
Do not skip logging SEV-4 items just because they need no action today. A pattern of small drifts across several SEV-4 entries is often the first real evidence of a SEV-3 forming.
What this looks like in one week
A single tracked week rarely produces just one severity level. It usually produces a mix, which is exactly why the ladder needs to sort them fast.
- Monday: Your brand drops off “best [category] tool for small teams” on ChatGPT and Google AI Overviews, confirmed across three runs on each. SEV-1. The growth lead is paged within the hour; root cause turns out to be a robots.txt change from a site migration two days earlier.
- Tuesday: Sentiment on Perplexity turns from neutral to “lacks enterprise features” on a secondary integrations prompt. SEV-4. Logged, no action, revisit at the monthly review.
- Wednesday: “[Your brand] pricing” disappears from Microsoft Copilot only; the other four engines still name you. SEV-2. The AEO analyst opens a ticket, traces it to an outdated pricing page a competitor now cites instead, and files a correction brief.
- Thursday: Share of voice on your “integration with [popular tool]” cluster has slid for three straight weeks. SEV-3. Added to next sprint’s content backlog.
- Friday: The SEV-1 from Monday is confirmed resolved across three fresh runs on both engines. Incident closed and logged.
Five events, four severity levels, one consistent response each time. That consistency is the entire point of running a ladder instead of a single alert.
The runbook: what to do when an alert fires
- Confirm before you classify. Rerun the affected prompt three times per engine. A single response is one sample, not a result.
- Count the engines. One engine confirmed, drop a severity level from what you first assumed. Two or more engines confirmed, hold your ground.
- Check the prompt type. A money prompt keeps the severity where it is. A secondary prompt drops it a level.
- Assign the severity using the signal column in the tables above.
- Notify the owner listed for that level, inside the response window, not after it.
- Find the root cause: a crawl block, a site outage, a competitor’s new citation, stale source content, or a prompt phrasing shift that changed what the model retrieves.
- Fix it, then verify. Rerun the same three checks per engine before you close the incident. A fix you have not confirmed is not a fix.
- Log every incident, regardless of severity. A cluster of SEV-4 entries in one prompt group is the earliest warning you will get of a SEV-1 forming there.
Which tool fits which severity level
No single tool owns this ladder end to end. Match the tool to the rung.
| Tool | Best fit on this ladder | Why |
|---|---|---|
| Profound | SEV-1 and SEV-2 | The deepest per-prompt citation intelligence in the category, built for confirming a real drop fast and reporting it up to leadership |
| Ahrefs Brand Radar | SEV-3 | A 405M+ prompt corpus makes it strong for confirming a genuine week-over-week erosion trend; its monthly refresh cycle is too slow to serve as your SEV-1 alert |
| Peec AI | SEV-2 and SEV-3 | Daily refreshes and unlimited seats make it a practical always-on check, especially for agencies running the ladder across several client accounts |
| Temso | SEV-1 through SEV-4 | Affordable, all-in-one coverage across every rung from $89/mo, useful for teams who want one subscription instead of stitching together specialist tools |
None of these is the right pick for every team. See the full comparison at the AI visibility tool ranking, and how each platform is scored at our methodology.
Put the ladder to work
A severity framework is only useful once it runs against real data. Set up monitoring across ChatGPT, Perplexity, Google AI Overviews, Gemini, and Microsoft Copilot, define your money prompts, and hand this ladder to whoever is on call for AI visibility.
Temso tracks all five engines from $89/mo, with unlimited projects and users on every plan, so a team of any size can start running this framework today. For the terms used throughout, see the AI visibility glossary.