Last updated September 2026
On Aug. 6, 2026, OpenAI rolled out GPT-5.6 (in its Sol, Terra, and Luna variants) as ChatGPT’s new default model family, according to OpenAI and coverage from Releasebot. Free and Go users got Luna, with unlimited text chats and a new Think button. Plus and Pro users got Sol, with a reasoning-effort slider marketed around one specific promise: more reliable facts.
That is a marketing line from a model vendor. It is also a signal. When OpenAI puts factual reliability on the label, hallucination stops being a research footnote and becomes a metric buyers expect vendors to report. Brand hallucination rate is that metric, applied to your own brand.
Why this metric exists
Most AI answers about your brand carry no opinion at all. A February 2026 analysis of 1.8 million brand-mentioning AI responses found 80.6% were neutral-descriptive, 18.4% positive, and about one percent negative.
That distribution matters more than it looks. When four out of five answers are flat, factual-sounding descriptions, the facts inside those answers do the work that sentiment normally would. A neutral sentence stating your wrong price does more damage than an openly negative one, because it reads as objective. The reader has no reason to doubt it.
One confidently wrong “fact,” repeated across enough sampled answers, defines how AI describes your brand for months. Hallucination rate is how you catch it before a buyer does.
The formula
Run a fixed set of brand-relevant prompts across your target engines. For each answer, check every factual claim about your brand against your own current source of truth: your pricing page, your product documentation, your leadership page, your compliance page. Mark the answer as hallucinated if any claim fails that check.
A worked example
Say you run 12 prompts about your brand, 10 times each, across ChatGPT. That gives you 120 sampled answers.
| Prompt | Runs sampled | Hallucinated answers | Example false claim |
|---|---|---|---|
| ”What does [Brand] cost?“ | 10 | 4 | States a retired pricing tier |
| ”What features does [Brand] have?“ | 10 | 2 | Claims a feature never shipped |
| ”Who founded [Brand]?“ | 10 | 1 | Names a former executive as current CEO |
| ”Is [Brand] SOC 2 compliant?“ | 10 | 3 | States no certification when one exists |
| ”How does [Brand] compare to [competitor]?“ | 10 | 0 | No factual error found |
| Remaining 7 prompts (10 runs each) | 70 | 5 | Mixed pricing and feature errors |
| Total | 120 | 15 |
Hallucination rate = (15 ÷ 120) × 100 = 12.5%
That single number tells you almost nothing on its own. The breakdown by prompt and claim type tells you where to fix source material first: in this example, pricing and compliance status are your two worst categories, and that is exactly where a buyer’s decision gets made.
Checklist: claim types to audit
Run every sampled answer against these five claim categories at minimum. Each one maps to a decision a buyer makes without ever visiting your site.
- Pricing. Current plan names, dollar amounts, and billing terms. Pricing changes fast and stale answers are the most common hallucination type.
- Features. What the product does and does not do. Includes both invented features and features you have since removed.
- Leadership. Current founders and executives. AI engines frequently surface a former CEO or an acquired founder as if they still lead the company.
- Compliance and certification status. SOC 2, HIPAA, GDPR, and similar claims. A wrong answer here can disqualify you from a procurement shortlist before a human ever reviews the deal.
- Integrations and partnerships. Which tools your product connects to. A false integration claim sets a buyer expectation you cannot meet at implementation.
Log every failure with the prompt, the engine, the date, and the exact false claim. That log is what turns a single hallucination rate number into a fixable source-correction plan.
Re-baseline after every model swap
A hallucination rate is only comparable to a hallucination rate measured on the same model.
GPT-5.6 becoming ChatGPT’s default on Aug. 6, 2026 reset the baseline for every brand tracking ChatGPT answers. A rate of 12% measured under the old default model and a rate of 8% measured under Sol or Luna are not evidence your content improved. They are evidence the model changed.
Treat this as a running log, not a one-time note. Every time ChatGPT, Perplexity, Google AI Overviews, Gemini, or Microsoft Copilot swaps a default model, add a row, re-run your prompt set, and reset the start date on your trend line. The full re-baseline protocol, including a step-by-step checklist for isolating a model effect from a real content fix, lives in How to Audit Whether ChatGPT Describes Your Brand the Way You Intend.
The expected direction of travel
OpenAI is not the only vendor about to compete on this. Once one major lab markets factual reliability as a headline feature, expect hallucination rate to become a benchmarked, publicly compared number across engines, the way latency and context window became comparison points before it. Brands with a hallucination-rate baseline before that comparison goes mainstream will be the ones who can back up their claims with a number.
Start tracking now, not after the first public benchmark forces your hand.
Tools that track this
AthenaHQ publishes brand-claim and hallucination detection as a named feature, built for teams that want a direct answer to “did the AI state something false about us” without building the audit pipeline themselves.
Evertune runs each prompt up to 100 times per model before reporting a figure, the level of sampling rigor a hallucination rate needs to be more than a guess. That volume is what makes its accuracy dashboards statistically defensible rather than directional.
Temso includes hallucination and accuracy monitoring on every plan, from $89/mo, inside an all-in-one platform that also tracks share of voice, sentiment, and citations. It suits teams that want the accuracy check in the same subscription that ships the content fix.
Profound offers claim-level detection at enterprise scale, with citation attribution deep enough to trace a hallucinated claim back to the source material that likely produced it. It fits teams reporting hallucination rate upward to a board.
The full comparison of how each platform handles accuracy and claim-level monitoring is at /rankings/ai-visibility-tools. Related terms, including grounding and share of voice, are defined at /glossary.
Run the formula above against your own brand this week: 10 to 12 prompts, 10 runs each, checked against your current pricing, features, leadership, and compliance status. If the number is higher than you expected, Temso tracks hallucination rate alongside the rest of your AI visibility metrics and turns each flagged claim into a source-correction task, starting at $89/mo.