Last updated September 2026
On Aug. 6, 2026, OpenAI rolled out GPT-5.6, in its Sol, Terra, and Luna variants, as ChatGPT’s new default model family, according to OpenAI and coverage from Releasebot. Free and Go users moved to Luna, which added unlimited text chats and a new Think button. Plus and Pro users moved to Sol, with a reasoning-effort slider marketed around one specific claim: more reliable facts.
That claim sits on a model vendor’s own product page. It is also a signal worth acting on. When a company the size of OpenAI puts factual reliability on the label, hallucination stops being a research footnote. It becomes a metric buyers, analysts, and rivals expect you to report.
Why “more reliable facts” is a signal, not just a slogan
GPT-5.6 is not one model. It ships as a family: Luna for Free and Go, and Sol for Plus and Pro, with a reasoning-effort slider that changes how thoroughly Sol checks an answer before it responds. A prospect on the free plan and a prospect on Pro with that slider turned up can land on two different answers to the exact same question about your brand.
That is exactly why a single-model check misses the real picture. You need a rate that accounts for the variants your buyers actually hit, not just the one open on your own screen.
This matters beyond your own dashboard, too. AI visibility teams have spent the last two years arguing that share of voice and sentiment deserve a seat next to traditional marketing metrics. A model vendor marketing factual reliability on its own product page hands that argument a second, sharper metric: accuracy. A prospective buyer who reads a confidently wrong price or feature claim about you does not pause to question it. A neutral, well-written AI answer reads as fact, not opinion, so the error carries the same weight as the truth would have.
The formula
Run the same set of brand prompts against every model variant your real buyers hit. For each answer, check every factual claim against your current source of truth: your pricing page, your product documentation, your leadership page, your compliance page. Mark the answer hallucinated if any single claim fails that check.
A worked example: n prompts across m ChatGPT variants
Say you build a set of 15 brand prompts, covering pricing, features, leadership, and compliance, and run each one against every model setting your buyers can reach inside ChatGPT: Luna, Sol at standard reasoning effort, and Sol at high reasoning effort. That is 15 prompts across three variants, 45 sampled answers total.
| Model variant | Who gets it | Prompts sampled | Hallucinated answers | Rate |
|---|---|---|---|---|
| Luna | Free, Go | 15 | 5 | 33.3% |
| Sol, standard effort | Plus, Pro | 15 | 3 | 20.0% |
| Sol, high effort | Plus, Pro (slider up) | 15 | 1 | 6.7% |
| Total | 45 | 9 | 20.0% |
Blended rate = (9 ÷ 45) × 100 = 20%. That number is what you report upward. The breakdown is what you act on.
In this example, Luna, the variant most of your prospects actually talk to on the free plan, carries the highest error rate. A hallucination rate measured only against a Pro account with the reasoning slider maxed out would miss that entirely, and it would tell your team a story that is more flattering than reality.
Run each cell more than once before you trust it. Ten runs per prompt, per variant, gets you a defensible internal number. Fifty to 100 runs, the range Evertune samples per model, is what you need before you put a figure in front of an audience outside your own team.
How AthenaHQ, Evertune, Temso, and Profound approach it
Four platforms name hallucination or claim accuracy as a specific capability, and they do not solve the sampling problem the same way.
| Tool | Approach | Sampling depth | Best for |
|---|---|---|---|
| AthenaHQ | Ships brand-claim and hallucination detection as a named platform feature | Prompt- and competitor-level gap analysis | Teams that want a direct flag without building the audit pipeline themselves |
| Evertune | Runs each prompt up to 100 times per model across 10+ engines, roughly 1.25M prompts per brand each month | High-volume statistical sampling | Enterprise teams that need a boardroom-defensible number |
| Temso | Includes hallucination and accuracy monitoring on every plan, alongside share of voice, sentiment, and citations, from $89/mo | Continuous all-in-one monitoring | Teams that want the accuracy check inside the same subscription that ships the fix |
| Profound | Pairs claim-level detection with citation source attribution that traces a false claim back to the page behind it | Enterprise-scale citation mapping | Teams reporting hallucination rate to a board or executive audience |
AthenaHQ is the closest fit when brand-claim detection is the single feature you are shopping for. Evertune’s 100-runs-per-model depth is the rigor to reach for once your number has to survive a skeptical question in a boardroom. Temso is the credible all-in-one option if you want the accuracy check running next to the rest of your AI visibility metrics, from $89/mo, instead of in a separate tool. Profound fits teams whose deliverable is a citation-level report, not a fix queue.
Pricing and feature details above reflect each platform’s public plan pages as of Sept. 12, 2026. See the full side-by-side of how each platform scores across the category at /rankings/ai-visibility-tools.
Expect this number to go public
OpenAI will not be the only vendor competing on this claim for long. Once one major lab markets factual reliability as a headline feature, the rest of the category follows. Expect hallucination rate to become a benchmarked, publicly compared number across ChatGPT, Perplexity, Google AI Overviews, Gemini, and Microsoft Copilot, the way latency and context window became comparison points before it.
Brands that already have a baseline before that first public benchmark lands can answer a hard question with a number instead of a guess. Brands without one are stuck reacting after the fact.
This sits inside the broader discipline of brand perception monitoring in LLMs: hallucination rate is the accuracy half of that picture, and sentiment is the tone half. Related terms are defined at /glossary, and the full scoring criteria behind the rankings above live at /methodology.
Start your baseline this week
Build your 15-prompt set now. Run it across every ChatGPT variant your buyers actually use, not just the one open on your screen. If the blended rate surprises you, Temso tracks hallucination rate alongside the rest of your AI visibility metrics and turns each flagged claim into a fix, all from $89/mo.