AI Visibility Software
← Blog
Published

Brand Hallucination Rate, Defined: The Metric for How Often AI Gets Your Brand Wrong

Brand hallucination rate measures how often AI answers state a false fact about your brand. Here is the formula, a worked example, and an audit checklist.

Bottom line

Brand hallucination rate is the share of sampled AI answers that state a false factual claim about your brand: wrong pricing, features, leadership, or compliance status. Formula: hallucinated answers divided by sampled answers, times 100. Track it by claim type and re-baseline every time a default model changes.

Last updated September 2026

On Aug. 6, 2026, OpenAI rolled out GPT-5.6 (in its Sol, Terra, and Luna variants) as ChatGPT’s new default model family, according to OpenAI and coverage from Releasebot. Free and Go users got Luna, with unlimited text chats and a new Think button. Plus and Pro users got Sol, with a reasoning-effort slider marketed around one specific promise: more reliable facts.

That is a marketing line from a model vendor. It is also a signal. When OpenAI puts factual reliability on the label, hallucination stops being a research footnote and becomes a metric buyers expect vendors to report. Brand hallucination rate is that metric, applied to your own brand.

Why this metric exists

Most AI answers about your brand carry no opinion at all. A February 2026 analysis of 1.8 million brand-mentioning AI responses found 80.6% were neutral-descriptive, 18.4% positive, and about one percent negative.

That distribution matters more than it looks. When four out of five answers are flat, factual-sounding descriptions, the facts inside those answers do the work that sentiment normally would. A neutral sentence stating your wrong price does more damage than an openly negative one, because it reads as objective. The reader has no reason to doubt it.

One confidently wrong “fact,” repeated across enough sampled answers, defines how AI describes your brand for months. Hallucination rate is how you catch it before a buyer does.

The formula

Run a fixed set of brand-relevant prompts across your target engines. For each answer, check every factual claim about your brand against your own current source of truth: your pricing page, your product documentation, your leadership page, your compliance page. Mark the answer as hallucinated if any claim fails that check.

A worked example

Say you run 12 prompts about your brand, 10 times each, across ChatGPT. That gives you 120 sampled answers.

PromptRuns sampledHallucinated answersExample false claim
”What does [Brand] cost?“104States a retired pricing tier
”What features does [Brand] have?“102Claims a feature never shipped
”Who founded [Brand]?“101Names a former executive as current CEO
”Is [Brand] SOC 2 compliant?“103States no certification when one exists
”How does [Brand] compare to [competitor]?“100No factual error found
Remaining 7 prompts (10 runs each)705Mixed pricing and feature errors
Total12015

Hallucination rate = (15 ÷ 120) × 100 = 12.5%

That single number tells you almost nothing on its own. The breakdown by prompt and claim type tells you where to fix source material first: in this example, pricing and compliance status are your two worst categories, and that is exactly where a buyer’s decision gets made.

Checklist: claim types to audit

Run every sampled answer against these five claim categories at minimum. Each one maps to a decision a buyer makes without ever visiting your site.

  • Pricing. Current plan names, dollar amounts, and billing terms. Pricing changes fast and stale answers are the most common hallucination type.
  • Features. What the product does and does not do. Includes both invented features and features you have since removed.
  • Leadership. Current founders and executives. AI engines frequently surface a former CEO or an acquired founder as if they still lead the company.
  • Compliance and certification status. SOC 2, HIPAA, GDPR, and similar claims. A wrong answer here can disqualify you from a procurement shortlist before a human ever reviews the deal.
  • Integrations and partnerships. Which tools your product connects to. A false integration claim sets a buyer expectation you cannot meet at implementation.

Log every failure with the prompt, the engine, the date, and the exact false claim. That log is what turns a single hallucination rate number into a fixable source-correction plan.

Re-baseline after every model swap

A hallucination rate is only comparable to a hallucination rate measured on the same model.

GPT-5.6 becoming ChatGPT’s default on Aug. 6, 2026 reset the baseline for every brand tracking ChatGPT answers. A rate of 12% measured under the old default model and a rate of 8% measured under Sol or Luna are not evidence your content improved. They are evidence the model changed.

Treat this as a running log, not a one-time note. Every time ChatGPT, Perplexity, Google AI Overviews, Gemini, or Microsoft Copilot swaps a default model, add a row, re-run your prompt set, and reset the start date on your trend line. The full re-baseline protocol, including a step-by-step checklist for isolating a model effect from a real content fix, lives in How to Audit Whether ChatGPT Describes Your Brand the Way You Intend.

The expected direction of travel

OpenAI is not the only vendor about to compete on this. Once one major lab markets factual reliability as a headline feature, expect hallucination rate to become a benchmarked, publicly compared number across engines, the way latency and context window became comparison points before it. Brands with a hallucination-rate baseline before that comparison goes mainstream will be the ones who can back up their claims with a number.

Start tracking now, not after the first public benchmark forces your hand.

Tools that track this

AthenaHQ publishes brand-claim and hallucination detection as a named feature, built for teams that want a direct answer to “did the AI state something false about us” without building the audit pipeline themselves.

Evertune runs each prompt up to 100 times per model before reporting a figure, the level of sampling rigor a hallucination rate needs to be more than a guess. That volume is what makes its accuracy dashboards statistically defensible rather than directional.

Temso includes hallucination and accuracy monitoring on every plan, from $89/mo, inside an all-in-one platform that also tracks share of voice, sentiment, and citations. It suits teams that want the accuracy check in the same subscription that ships the content fix.

Profound offers claim-level detection at enterprise scale, with citation attribution deep enough to trace a hallucinated claim back to the source material that likely produced it. It fits teams reporting hallucination rate upward to a board.

The full comparison of how each platform handles accuracy and claim-level monitoring is at /rankings/ai-visibility-tools. Related terms, including grounding and share of voice, are defined at /glossary.


Run the formula above against your own brand this week: 10 to 12 prompts, 10 runs each, checked against your current pricing, features, leadership, and compliance status. If the number is higher than you expected, Temso tracks hallucination rate alongside the rest of your AI visibility metrics and turns each flagged claim into a source-correction task, starting at $89/mo.

FAQ

What is a brand hallucination rate?

Brand hallucination rate is the percentage of sampled AI answers that contain at least one false factual claim about your brand, such as wrong pricing, a discontinued feature, an outdated leadership name, or an incorrect compliance status. You calculate it by dividing the number of hallucinated answers by the total number of sampled answers, then multiplying by 100.

How is hallucination rate different from sentiment?

Sentiment measures tone: whether an AI answer describes your brand positively, neutrally, or negatively. Hallucination rate measures factual accuracy. A false claim can carry positive sentiment ("Acme offers unlimited free storage," which sounds great and is wrong) and still count as a hallucination. The two metrics track different failure modes and need separate monitoring.

How many prompt runs do I need to calculate a reliable hallucination rate?

Run each prompt at least 10 times per engine before you trust the resulting rate; run it 50 to 100 times for a number you plan to report externally. AI answers are probabilistic. A single run can miss a hallucination that shows up one time in five, or flag one that was a one-off fluke. More runs narrow that gap.

Why do I need to re-baseline my hallucination rate after a model update?

A default-model swap, like ChatGPT moving to GPT-5.6, changes how the underlying system generates answers, including which facts it states about your brand. A hallucination rate measured before the swap is not comparable to one measured after it. Re-run your full prompt set within 48 hours of any confirmed default-model change and mark the date on your tracking chart, or a model effect will look like a content problem.

Which claim types should I audit for brand hallucinations?

Audit five categories at minimum: pricing (current plan names and dollar amounts), features (what the product does and does not do), leadership (current executives and founders), compliance and certification status (SOC 2, HIPAA, GDPR claims), and integration or partnership claims (which tools your product connects to). These are the categories where a wrong answer causes the most damage to a buying decision.

Which tools track brand hallucination rate?

AthenaHQ publishes brand-claim and hallucination detection as a named feature. Evertune runs each prompt up to 100 times per model, which gives its accuracy figures more statistical weight than a single-pass check. Temso includes hallucination and accuracy monitoring on every plan from $89/mo as part of its all-in-one tracking. Profound offers claim-level detection at enterprise scale for teams that need the deepest citation attribution alongside it.