Last updated September 2026. This is the initial publication of this checklist, tracking OpenAI’s Aug. 26, 2026 retirement of o3 in ChatGPT. Every future default-model swap gets a new row in the timeline table below.
What changed on Aug. 26, 2026
OpenAI shut off o3 in ChatGPT on Aug. 26, 2026, according to the OpenAI Help Center’s model release notes. That date closed a 90-day sunset window that opened on May 28, 2026, when OpenAI first flagged o3 for retirement.
GPT-5.6, across its Sol, Terra, and Luna variants, is now the only default model lineup in ChatGPT. It runs on Free, Go, Plus, and Pro. There is no legacy o3 toggle left to fall back on.
If your team runs prompt tests or citation audits against “ChatGPT,” the model answering those prompts just changed. That matters more than most marketers assume.
Why a model swap breaks your visibility trendline
Benchmark stability is a myth. A share-of-voice chart compares two points in time, and that comparison only holds if the thing generating the answer stayed the same between those points.
It didn’t. A different model can phrase things differently, pull from different sources, and mention different brands, even when nothing about your content, your citations, or your competitors changed. Treat the pre-swap and post-swap periods as one continuous line, and you will credit or blame your own GEO work for a shift the model caused on its own.
This is the same discipline covered in what prompt monitoring is and how it works: a prompt library, a schedule, a mention parser, and a delta between runs. A default-model swap is the one event that breaks the delta step. Everything before the swap and everything after it belongs in two separate eras, not one trendline.
Annotate, re-run, or reset? Use this decision box
Not every OpenAI update deserves the same response. Match the trigger to the action.
The Aug. 26, 2026 swap is a full reset. o3 didn’t get a point update, it disappeared, and the entire default lineup changed with it.
The 10-step re-baselining checklist
Work through these in order. Skipping steps is how teams end up reporting a “content win” or a “content loss” that was actually just OpenAI shipping a new model.
- Confirm the swap with an official source. Check the OpenAI Help Center’s model release notes, not a forum post or a screenshot. “ChatGPT feels different” is not confirmation. A dated changelog entry is.
- Log the exact date. Record the swap date, the outgoing model, and the incoming model in your tracking sheet or your monitoring tool’s annotation feature. Every future chart needs this line to reference.
- Freeze your pre-swap baseline. Snapshot the last clean run before the swap date and label it clearly. Don’t let new data quietly blend into an old trendline.
- Flag any in-flight runs. A scheduled run that landed inside the transition window, the 90 days between the sunset announcement and the cutover in this case, is transition data, not trend data. Mark it. Don’t delete it.
- Re-run your full tracked-prompt set on demand. Don’t wait for the next nightly cycle. Trigger an immediate run against the new default so you have a like-for-like comparison point as close to the swap date as possible.
- Compare mention rate, citations, and sentiment side by side. Line up the frozen baseline against the fresh run, prompt family by prompt family. Look for consistent movement across many prompts, not a single outlier.
- Separate model change from content change. If a prompt’s result moved and you touched no content, no citations, and no pages tied to that prompt, the model caused the shift. Not your GEO work.
- Reset your trend window. Start a new multi-run baseline period, five to seven days minimum, before you trust week-over-week deltas again. One post-swap run is a sample, not a trend.
- Re-date your reporting. Add a note to any dashboard, deck, or report a stakeholder will see. Show the pre-swap and post-swap eras as separate segments, not one continuous line.
- Add the swap to a permanent changelog. Keep a running, dated log of every default-model swap your team has re-baselined against. Start with the timeline below. The next swap gets easier to catch once you have a place to log it.
The 2026 ChatGPT default-model swap timeline
This table is the running, dated record. Every future swap that forces a re-baseline adds a new row.
| Date | What changed | Source |
|---|---|---|
| Aug. 7, 2025 | GPT-5 replaces GPT-4o as the ChatGPT default across paid tiers. | OpenAI, August 2025 |
| Nov. 12, 2025 | GPT-5.1 becomes the new ChatGPT default. | OpenAI, November 2025 |
| May 28, 2026 | OpenAI flags o3 for retirement and opens a 90-day sunset window. | OpenAI Help Center |
| Aug. 26, 2026 | o3 is retired from ChatGPT; GPT-5.6 (Sol, Terra, and Luna) becomes the full default lineup across Free, Go, Plus, and Pro. | OpenAI Help Center, Model Release Notes |
OpenAI shipped several point releases between GPT-5.1 and GPT-5.6 across late 2025 and into 2026. For the exact date of a specific point release, check the OpenAI Help Center’s model release notes directly. This table tracks default-lineup swaps significant enough to justify a re-baseline, not every point release.
Which tools support annotate-and-re-run workflows
Feature availability below reflects each tool’s public plan pages as of Sept. 1, 2026.
Peec AI tracks prompts daily and lets you drop a note on a specific date in the dashboard, so a swap like this one gets pinned to the exact day your numbers moved. It also supports triggering a fresh prompt run outside its normal schedule, which is exactly what step five above calls for.
Profound timestamps every citation pull inside its citation maps, giving you a clean marker for “last clean run before Aug. 26, 2026” without extra spreadsheet work. Its Prompt Volumes view also makes it easier to separate a swap-driven shift from a genuine change in what buyers are asking.
Temso is a credible all-in-one option here: annotation and on-demand prompt re-runs live inside the same $89/mo subscription that also tracks mentions, citations, and sentiment, so re-baselining after a swap like this one doesn’t require a second tool or a manual export.
Otterly.AI is worth including for teams on a tighter budget. Its prompt-level tracking supports the same freeze-and-compare pattern, even though its annotation options are lighter than what Peec AI or Profound offer.
For a full side-by-side of the category, see the AI visibility tool rankings. If terms like share of voice or citation rate are new to your team, the AI visibility glossary defines them in plain language. This checklist follows the same sampling principle laid out in our methodology: never trust a single run, and always compare like periods to like periods.
Start your re-baseline now
Don’t wait for next week’s dashboard to look strange before you act. Freeze today’s data, log Aug. 26, 2026 as the swap date, and re-run your tracked prompts against GPT-5.6 this week.
If your current tool can’t annotate a swap or trigger an on-demand re-run, that gap will cost you again the next time OpenAI ships a new default, and it will happen again. Temso covers both in one $89/mo subscription. Start a free trial and re-baseline your first prompt set today.