When OpenAI replaced GPT-4o on 7 August 2025, the story became "people mourning their AI companion." A 9,908-post Reddit corpus tells a different one: capability complaints led on launch day, and the phrase "my long threads broke" hides at least four separate causes — one of them documented a month before GPT-5 shipped.
Numbers marked Analysis are this project's classification of a Reddit corpus, not a published statistic. Numbers and events marked Source are drawn from published studies, OpenAI documentation, or filed bug reports and were verified against the original on 27 Aug 2026.
Users reported a single experience — the model losing track of a long conversation. The evidence shows that experience had at least four distinct roots, colliding in the same fortnight. No user could tell them apart from inside a chat. Source
A bug report on OpenAI's own developer forum — 8 July 2025, a month before GPT-5 — already described the symptom: after roughly 30–50 messages, GPT-4o forgot the earlier part of the same chat. Within-thread forgetting existed on 4o before GPT-5 shipped.
A structured report on OpenAI's cookbook issue tracker claims GPT-5 loses previously supplied context across turns where GPT-4 did not — the reporter explicitly calls it a regression. So it is not only baseline: a technical user who could compare claimed a genuine GPT-5-specific loss.
A separate forum thread describes long conversations progressively breaking: parts of the dialogue disappearing, the interface returning a content-loading failure. A frontend/persistence bug — messages actually vanishing — that looks identical to the user but has a completely different root cause.
OpenAI's launch auto-routed retired-model chats to the nearest GPT-5 variant; "Auto" itself routes across internal models. In the corpus, users felt this as sudden tonal or capability shifts mid-conversation and theorised a hidden "safety router." The routing was real; the mechanism they inferred was mostly unverifiable.
Four causes, one symptom, same two weeks, none separable from inside a chat. The existing literature never untangled this because it never examined the continuity axis at all.
Posts carrying a capability signal vs an emotional signal, by day. On 7 August — launch — task complaints outnumbered emotional ones; emotion overtook on the 8th. Analysis
Share of all 9,908 posts carrying each signal (single-label by presence). Emotional leads capability, but only just — and most posts carry neither. Analysis
Total upvotes on posts carrying each complaint type. The companion story dominated the media; in the discourse, capability and long-thread complaints sit at the same order of magnitude as tone. Higher = more community weight. Analysis
Of 750 posts matching memory/context keywords, hand-reading splits them into four distinct phenomena — the same conflation as the causes above, visible in the data. Only the first is the within-conversation continuity question. Analysis
"GPT-5 hallucinates. It forgets context. It feels unstable and inconsistent. GPT-4o was sharp, focused, and reliable."
"…it can't follow complex situations as well, doesn't reason as well within the chat."
A physics researcher reported GPT-5 "paralyzed my research" across four separate professional uses.
The 264 come from coders, novelists, and researchers independently — not the companion subreddit. In the comments, corroboration of the continuity complaint clearly outweighed organised "it's placebo" skepticism, though archived comment scores are unreliable so no precise ratio is claimed. The sharpest tell: users who switched back to legacy GPT-4.1 reported continuity returned — same person, same threads, swap the model, symptom changes.
The published side and the undocumented side of the same question. The gap is the story — users were left to reverse-engineer. Source
OpenAI states these by plan; users documented them and raged at the Plus ceiling.
What happens when a thread exceeds the window is undocumented.
Every specific "it auto-compresses" claim online is inference from behaviour, not OpenAI documentation. With Plus at ~32K, heavy users' threads were already shedding early turns before GPT-5 arrived — Cause A, restated from the platform side. The "Context Degradation Syndrome" term circulating in blogs has no official standing and is excluded here as a source.
Every prior study of this event measured attachment, grief, or governance. The right-hand column is the gap this page fills. Click a header to sort. Source · all verified 27 Aug 2026
| Study ▼ | Year ▼ | Corpus ▼ | Platform ▼ | Focus ▼ | Capability / continuity? ▼ |
|---|---|---|---|---|---|
| Keeman & Keeman | 2026 | 2,100 | Scored responses | Empathy / safety | Empathy indistinguishable across models (p=0.115); safety posture shifted. No |
| Wu | 2026 | 61,846 | X | Themes / governance | Widest net; version succession ≠ effective replacement on the user side. No |
| Lai | 2026 | 1,482 | X (#keep4o) | Resistance / rights | Instrumental vs relational attachment; loss of choice drove protest. No |
| Naito | 2025 | 150 | Multi-platform | Attachment / culture | Attachment framing 78% JP vs 38% EN (χ²=24.9). No |
| Madsen | 2025 | conceptual | Discourse | Parasocial theory | Frames it as a parasocial breakup / breach of an "affective contract". No |
| Donard & Ribeiro | 2026 | 3,668 | Reddit + YouTube | Mourning / MH | Nine themes; attachment (425), disappointment (263), grief (110). No |
| This page | 2026 | 9,908 | Capability / continuity | The unfilled axis: within-thread continuity, capability-led timeline, four-cause separation. Yes |
Because Wu already covered X at 61,846 posts, Reddit is the differentiator — it is where workflow and long-thread complaints live rather than hashtag activism.
#keep4o spans two distinct moments. This page's corpus and most studies cover the first; one study (Keeman) is framed around the second. Source
Corpus. 9,908 unique Reddit posts + 5,270 comments from r/ChatGPT, r/ChatGPTPro, r/MyBoyfriendIsAI and r/OpenAI, 6 Aug–30 Sep 2025, retrieved via the Arctic Shift archive. Post-level upvote scores are intact; archived comment scores are not, so comments are sampled thematically rather than ranked.
Classification is keyword-based, hand-verified only on the 264-post continuity bucket and every quoted item. The macro split and upvote-weighted figures describe the shape and proportion of the discourse — they are not measurements of the model, and "forgets after two messages" is user perception, likely hyperbole. Complaint-cause theory counts carry more false positives than the topic buckets.
Two archive traps were handled (and are worth recording): the API caps at 100 results per call, so wide date ranges silently return only the newest slice — every launch-week day exceeded the cap and needed recursive sub-day slicing (r/ChatGPT, Aug 8: 100 → 620 posts). And a 422 on high-volume windows is a server hiccup, not an empty result — the same window returns 50 rows at a smaller limit, so naive handling under-reports the busiest slices.
Sourcing. Primary/official only — arXiv, OpenAI documentation, filed forum and issue reports. Forum/issue items are individual reports, not audited by OpenAI, and are treated as credible technical testimony rather than established fact. Corpus is Reddit-only; X is covered by Wu.