AI Visibility Software
← Blog
Published

The AI Visibility Incident Ladder: A SEV-1 to SEV-4 Severity Framework for Brand Visibility Drops

A SEV-1 to SEV-4 severity framework for AI visibility incidents, with detection signals, response windows, owners, and a runbook for a brand visibility drop.

Bottom line

The AI Visibility Incident Ladder classifies brand-visibility drops into four severity levels: SEV-1 (gone from a money prompt, two or more engines, respond within hours) down to SEV-4 (single-engine sentiment drift, log and monitor). Each level defines a signal, a response window, an owner, and an action.

Last updated September 2026

A brand disappears from a ChatGPT answer on a Tuesday morning. Someone notices, drops a screenshot in Slack, and the channel treats it like a five-alarm fire. Half the time, that reaction is correct. The other half, it is one probabilistic response among many, and the brand shows up fine on the next run.

Most teams have no way to tell which situation they are in. They have a monitoring dashboard, a gut feeling, and no defined line between “this is urgent” and “this can wait.” That gap is expensive. Real incidents sit unaddressed because they look like noise, and noise gets escalated because nobody set a bar in advance.

Site reliability teams solved this exact problem decades ago. A SEV-1 outage pages an engineer at 3 a.m. A SEV-4 bug goes in next sprint’s backlog. Nobody argues about which is which, because the criteria are written down before the pager ever goes off. AI visibility needs the same discipline. This ladder gives you four severity levels, each with a detection signal, a response window, an owner, and an action, so “our visibility dropped” stops being a feeling and starts being a classification.

The stakes are real. More than half of B2B software buyers, 51%, now begin their research inside an AI chatbot rather than a traditional search engine (G2, “The Answer Economy,” April 2026). A SEV-1 sitting open for a week is not a dashboard problem. It is buyers seeing a competitor on the exact query where you should have shown up first.

Why cross-engine confirmation matters before you panic

AI engines do not agree with each other nearly as often as most teams assume. BrightEdge’s AI Catalyst research found that brand mentions disagreed 61.9% of the time across Google AI Overviews, AI Mode, and ChatGPT, the three engines it tracked; only 33.5% of queries returned the same brand names across all three (BrightEdge AI Catalyst research, July 2025).

That fact should change how you triage. A brand missing from one engine, on one run, is close to the normal state of the system. A brand missing from the same money prompt on two or more engines, confirmed across repeated runs, is a different thing entirely. It is the disagreement rate collapsing into agreement, and agreement in this data is rare enough to matter.

This is the logic behind every threshold on the ladder below. The higher the severity, the more engines and the more repeated runs it takes to confirm the drop is real.

The ladder at a glance

Four rungs, ranked by how much revenue-relevant visibility is actually at risk.

SEV-1  CRITICAL   Page someone now      Money prompt gone, 2+ engines
SEV-2  HIGH       Respond same day      Money prompt gone, 1 engine; or sentiment turns negative, 2+ engines
SEV-3  MODERATE   Fix this week         Share of voice erodes on a tracked cluster; or one fact is wrong
SEV-4  LOW        Log and monitor       Sentiment drifts on one engine, secondary prompt only

A money prompt is a high-intent query, the kind closest to a buying decision: “best [category] tool,” “[brand] pricing,” “[brand A] vs. [brand B].” A secondary prompt is everything else: awareness-stage questions, feature explainers, general category education. The same drop matters far more on a money prompt than on a secondary one, which is why prompt type sets the ceiling on severity before engine count even comes into play.

How to define your money prompts before you need this ladder

The ladder only works if your money-prompt list exists before an incident hits. Building it takes an afternoon, not a quarter.

  1. Pull the last 10 to 20 questions your sales team says buyers actually asked during a demo or a win-loss call. Use their words, not your marketing copy.
  2. Add the two or three comparison prompts buyers run most: “[your brand] vs. [top competitor]” and “best [category] for [your core segment].”
  3. Cap the list at 10 to 15 prompts. A list you cannot check daily is a list nobody checks.
  4. Review the list every quarter. Buyer language shifts, and a stale prompt list will miss a real SEV-1 because you stopped tracking the query that mattered.

Everything else on this page, the tables, the runbook, the tool fit, assumes this list already exists. Skip it and every severity call downstream becomes a guess.

SEV-1: Critical

Your brand is gone from a money prompt on two or more of ChatGPT, Perplexity, Google AI Overviews, Gemini, or Microsoft Copilot, and the drop holds across repeated runs. This is the only level where you page someone the same hour you see it.

SignalResponse windowOwnerAction
Brand absent from a money prompt on 2+ engines, confirmed across 3 runs per engineAcknowledge in 1 hour; root cause in 4 hoursGrowth or marketing lead, plus AEO analystRerun the prompt, rule out a site outage or crawl block, check for a competitor’s new citation-worthy asset, escalate same day

A SEV-1 that sits unaddressed for a week is not an anomaly anymore. It is your new baseline, and every day it stays open is a day buyers see a competitor instead of you on a query you should own.

SEV-2: High

Either your brand is gone from a money prompt on one engine only, or sentiment turns negative on two or more engines for a prompt that still names you. Both need a response inside the day, but neither needs a 3 a.m. page.

SignalResponse windowOwnerAction
Money prompt gone on 1 engine; or sentiment flips negative on 2+ engines, confirmed across 3 runsAcknowledge in 24 hours; resolve in 3 business daysAEO analyst or content leadDiagnose the single-engine cause, draft a correction brief, check daily until it clears

A single-engine dropout at SEV-2 is often the leading edge of a SEV-1. Treat it as a warning, not a coincidence, and check the other engines again before you close it.

SEV-3: Moderate

Share of voice on a tracked prompt cluster has been declining for two or more consecutive weeks, or one factual inaccuracy (an old price, a discontinued feature) has appeared on a single engine with no urgent revenue exposure attached to it.

SignalResponse windowOwnerAction
Sustained week-over-week share-of-voice decline on a prompt cluster; or a factual error on 1 engineLog within the week; fix in the next content sprintContent or SEO teamAdd to the backlog, correct the source content, request a recrawl, note it in the weekly review

Most AI visibility work happens at this level. It is not urgent, but it compounds if nobody works the backlog.

SEV-4: Low

Sentiment drifts slightly on a single engine, on a prompt that is not a money prompt, or a competitor gains modest ground on a low-intent query. Nothing here needs a response today.

SignalResponse windowOwnerAction
Minor sentiment drift, 1 engine, secondary prompt onlyLog and monitor; revisit at the next monthly reviewWhoever owns the monitoring dashboardNote it in the tracker; act only if it persists past a month

Do not skip logging SEV-4 items just because they need no action today. A pattern of small drifts across several SEV-4 entries is often the first real evidence of a SEV-3 forming.

What this looks like in one week

A single tracked week rarely produces just one severity level. It usually produces a mix, which is exactly why the ladder needs to sort them fast.

  • Monday: Your brand drops off “best [category] tool for small teams” on ChatGPT and Google AI Overviews, confirmed across three runs on each. SEV-1. The growth lead is paged within the hour; root cause turns out to be a robots.txt change from a site migration two days earlier.
  • Tuesday: Sentiment on Perplexity turns from neutral to “lacks enterprise features” on a secondary integrations prompt. SEV-4. Logged, no action, revisit at the monthly review.
  • Wednesday: “[Your brand] pricing” disappears from Microsoft Copilot only; the other four engines still name you. SEV-2. The AEO analyst opens a ticket, traces it to an outdated pricing page a competitor now cites instead, and files a correction brief.
  • Thursday: Share of voice on your “integration with [popular tool]” cluster has slid for three straight weeks. SEV-3. Added to next sprint’s content backlog.
  • Friday: The SEV-1 from Monday is confirmed resolved across three fresh runs on both engines. Incident closed and logged.

Five events, four severity levels, one consistent response each time. That consistency is the entire point of running a ladder instead of a single alert.

The runbook: what to do when an alert fires

  1. Confirm before you classify. Rerun the affected prompt three times per engine. A single response is one sample, not a result.
  2. Count the engines. One engine confirmed, drop a severity level from what you first assumed. Two or more engines confirmed, hold your ground.
  3. Check the prompt type. A money prompt keeps the severity where it is. A secondary prompt drops it a level.
  4. Assign the severity using the signal column in the tables above.
  5. Notify the owner listed for that level, inside the response window, not after it.
  6. Find the root cause: a crawl block, a site outage, a competitor’s new citation, stale source content, or a prompt phrasing shift that changed what the model retrieves.
  7. Fix it, then verify. Rerun the same three checks per engine before you close the incident. A fix you have not confirmed is not a fix.
  8. Log every incident, regardless of severity. A cluster of SEV-4 entries in one prompt group is the earliest warning you will get of a SEV-1 forming there.

Which tool fits which severity level

No single tool owns this ladder end to end. Match the tool to the rung.

ToolBest fit on this ladderWhy
ProfoundSEV-1 and SEV-2The deepest per-prompt citation intelligence in the category, built for confirming a real drop fast and reporting it up to leadership
Ahrefs Brand RadarSEV-3A 405M+ prompt corpus makes it strong for confirming a genuine week-over-week erosion trend; its monthly refresh cycle is too slow to serve as your SEV-1 alert
Peec AISEV-2 and SEV-3Daily refreshes and unlimited seats make it a practical always-on check, especially for agencies running the ladder across several client accounts
TemsoSEV-1 through SEV-4Affordable, all-in-one coverage across every rung from $89/mo, useful for teams who want one subscription instead of stitching together specialist tools

None of these is the right pick for every team. See the full comparison at the AI visibility tool ranking, and how each platform is scored at our methodology.

Put the ladder to work

A severity framework is only useful once it runs against real data. Set up monitoring across ChatGPT, Perplexity, Google AI Overviews, Gemini, and Microsoft Copilot, define your money prompts, and hand this ladder to whoever is on call for AI visibility.

Temso tracks all five engines from $89/mo, with unlimited projects and users on every plan, so a team of any size can start running this framework today. For the terms used throughout, see the AI visibility glossary.

FAQ

What is the AI Visibility Incident Ladder?

The AI Visibility Incident Ladder is a four-level severity framework for AI brand visibility incidents, adapted from the SEV-1 to SEV-4 scale that site reliability engineering teams use for system outages. It ranks a visibility drop by how much revenue-relevant exposure is at risk, from SEV-1 (a money prompt gone on two or more engines) to SEV-4 (sentiment drift on a single engine, secondary prompt only). Each level defines a detection signal, a response window, an owner, and a required action.

What counts as a SEV-1 AI visibility incident?

A SEV-1 incident is a brand disappearing from a money prompt, a high-intent query such as "best [category] tool" or "[brand] pricing", on two or more of ChatGPT, Perplexity, Google AI Overviews, Gemini, or Microsoft Copilot, confirmed across repeated runs. The two-engine threshold matters: BrightEdge's AI Catalyst research found that AI engines disagree on brand recommendations 61.9% of the time, so a single-engine dropout is often normal variation. A drop that holds across two or more engines is not noise.

How fast should you respond to a SEV-1 drop?

Acknowledge a SEV-1 within one hour and identify the root cause within four hours. Confirm the drop is real by rerunning the affected prompt three times per engine, rule out an obvious cause such as a site outage, a robots.txt block, or a competitor's newly cited asset, then escalate to a marketing or growth lead the same day. A SEV-1 left open for a week is not an anomaly. It is the new baseline.

Why does a brand disappear from ChatGPT answers?

The most common causes are a crawl or indexing block that keeps an AI bot from reading the site, a competitor publishing a more citable or more recent asset on the same query, stale source content the model still relies on, or a prompt phrasing shift that changes which sources the model retrieves. Confirming which cause applies is the first step of the incident runbook. The fix depends entirely on the cause.

Who should own AI visibility incident response?

Ownership should scale with severity. SEV-1 and SEV-2 incidents need a named owner, typically a growth or marketing lead plus whoever runs AEO monitoring, because a revenue-relevant prompt is at stake and the response window is measured in hours. SEV-3 and SEV-4 can sit with the content or SEO team as part of a normal weekly and monthly review, since the cost of a slower response is lower.

How is this different from a normal SEO ranking-drop alert?

A ranking-drop alert tracks one deterministic number: a position on a search results page. AI visibility is probabilistic, since the same prompt can return a different answer on back-to-back runs, and engines frequently disagree with each other on which brands to name. The Incident Ladder is built for that reality: it requires confirmation across multiple runs, and at the SEV-1 level across multiple engines, before a drop counts as a real incident. This filters out normal model variance without missing a genuine one.