Context rot is a reproducible drop in a large language model's accuracy on the exact same task, simply because the surrounding conversation has grown, even while the model is nowhere close to running out of room to keep reading. One widely cited study measured just how far that drop can go: a model handed the correct answer document scored 52.9 percent, worse than the 56.1 percent it scored with no documents in front of it at all. A ChatGPT or Claude conversation that started sharp an hour ago starts hedging, repeating a question it already answered, or flatly contradicting something it said ten messages back, and this is why: a chat can feel like it is getting dumber well before any token counter comes close to its limit. Nothing about the questions got harder. The conversation just got longer. Three distinct failures get blamed for this and lumped together as one thing, and a comparison table further down untangles which is which and what fixes each.
Running out of the context window is a different failure, and it is the one most people assume is happening when a chat starts sliding. The context window is a hard, visible wall: the model genuinely cannot see a message that has scrolled past its input limit, and it has to drop or summarize old turns to keep going. Context rot shows up before that wall, often long before it. A conversation sitting at a small fraction of its advertised token limit can already be producing worse answers than the same questions asked fresh, because length itself, not available room, is doing the damage.
Context rot is a recent name for the pattern: a commenter on Hacker News coined it in June 2025 to describe large language models getting less effective simply as their input grows, and the name stuck because engineers kept re-discovering the same pattern independently across completely different products.
Most explanations stop at naming the symptom. Fewer explain the actual mechanism in language that does not require a machine learning background. Two separate things happen inside a model every time a conversation grows by another message, and they compound each other. There is also a one-line test further down for telling, in your own chat, whether length is the real culprit before doing anything else.
Two Reasons a Long Chat Gets Worse With Room Left in the Window
Every new message splits a fixed attention budget
When a model generates its next word, it does not read the conversation the way a person does, skimming for the relevant part and ignoring the rest. It computes a relevance score between the word it is about to write and every single token currently sitting in its context, then normalizes those scores so they add up to a fixed total, a method called attention. Picture that total as a spotlight with a set amount of light to hand out. Early in a conversation, a few hundred tokens are splitting that light, so each one gets a usable share. Forty exchanges in, thousands of tokens, including every prior back-and-forth, tangent, and pasted paragraph, are all competing for a slice of that same fixed beam. The fact you need is still technically there. That fact is just getting a dimmer share of attention than it got when the conversation was short.
Anthropic's own engineering team describes this directly: a model must treat context as a finite resource with diminishing marginal returns, because it draws on what the team calls an attention budget, and every new token added depletes it. That framing comes from the company building the models, not an outside critic: the constraint is architectural, not a bug a future update quietly patches away.
Position math gets fuzzier the further it stretches
The second mechanism is less well known and rarely explained outside research papers. To keep track of word order, most modern language models tag each token's position using a technique called rotary position encoding, RoPE for short, which represents position as an angle of rotation rather than a plain counted index. Picture a clock hand: the model judges how far apart two tokens are partly by comparing the angle of their hands. Near the start of a conversation, those angle differences are small and clean, easy to read precisely. As the conversation stretches on, the hands have spun through far more rotations, and the fine angle differences the model relies on to tell '40,000 tokens apart' from '41,000 tokens apart' become harder to distinguish, because the model is now doing position math on ranges it rarely or never saw stretched this far during training.
Researchers studying this directly found that at long range, RoPE forces some of the dimensions inside a model's attention heads to rotate through such extreme angles that those dimensions stop being useful for telling positions apart, a problem they call dimension inefficiency. Positional degradation is a separate failure from attention dilution: attention dilution means a real fact gets a thinner slice of focus, while positional degradation means the model's internal sense of where that fact sits, and how far it is from the current question, gets blurrier the longer the conversation runs.
Your AI is not running out of room. It is running out of attention. More tokens dilute that attention, and longer range fuzzes the model's sense of position, both well before the context window is anywhere near full.
What the Research Found
The clearest early evidence came from a multi-document question-answering study out of Stanford, UC Berkeley, and Samaya AI. Researchers gave models a set of documents with exactly one containing the answer, then moved that answer document to different positions in the input. Accuracy formed a distinct U-shape: highest when the answer sat at the very start or the very end, and lowest in the middle. In the worst tested case, GPT-3.5-Turbo's accuracy dropped by more than 20 percentage points once the answer moved to the middle of the input, low enough that the model performed worse with the correct document sitting in front of it than it did with no documents at all, which scored 56.1 percent from memory alone. Handing the model the right answer in the wrong seat did worse than handing it nothing.
Lost in the middle, as that positional pattern is now widely called, is closely related to context rot but not identical to it. A separate report from Chroma ran 18 frontier models, spanning the GPT, Claude, Gemini, and Qwen families, on controlled tasks that held difficulty constant and varied only input length. Every single model degraded as input grew, including on a task as simple as copying a run of repeated words back verbatim. Length on its own broke performance, independent of where any particular fact happened to sit.
A more demanding version of the same test comes from Adobe Research and LMU Munich, published as NoLiMa, which replaced simple keyword matching with questions that required inference: given only that a character lives next to a specific landmark, the model had to work out which city that placed them in, rather than spotting a literal word match. At roughly 32,000 tokens of input, well inside every major model's advertised window, GPT-4o's accuracy on this task fell from 99.3 percent to 69.7 percent. Across the same test, Claude 3.5 Sonnet fell from 88 percent to 30 percent, Gemini 2.5 Flash from 94 percent to 48 percent, and Llama 4 Scout from 82 percent to 22 percent. The harder the retrieval task leaned on genuine reasoning rather than exact wording, the steeper the fall.
The study closest to an actual chat window comes from Microsoft Research and Salesforce Research. Instead of retrieval, the team measured what happens across real multi-turn conversations by taking fully specified tasks and releasing the instructions one fragment per turn instead of all at once, the way a real conversation with a human tends to unfold. Across more than 200,000 simulated conversations and six different generation tasks, every model tested performed worse in the multi-turn setting than the single-turn version of the identical task, by an average of 39 percent. Most of that loss was not information going missing. It came from models locking onto a wrong assumption in an early turn and building on it rather than recovering once later turns clarified the request, the exact pattern behind a chat that drifts off track without warning the longer it runs.
Three Failures People Lump Together
People talk about hitting the context window limit, lost in the middle, and context rot as one thing, but they are three distinct failures with three distinct fixes, and treating them as interchangeable is why they so often try the wrong fix first.
| What You're Seeing | Context Window Limit | Lost in the Middle | Context Rot |
|---|---|---|---|
| What triggers it | Total tokens exceed the model's hard input limit | The needed fact sits in the middle of a long prompt | The prompt's total length grows, whatever position the facts sit in |
| What it looks like | Old messages get dropped or silently summarized away | The model answers as if the fact were never mentioned | Answers get vaguer, more hedged, or quietly inconsistent |
| Does the window need to be full? | Yes, by definition | No | No, it starts well before the window fills |
| What helps | Start a fresh chat or restate the essentials | Move the key fact near the start or end of the prompt | Trim the conversation down to what is still relevant |
What Helps Mid-Chat
Waiting for a bigger context window fixes none of the three failures in the comparison above: it raises the ceiling on the hard wall, but it does nothing about attention getting diluted or position math getting fuzzier long before the model reaches that ceiling. The practical fixes are about managing what is inside the window, not the size of the window itself.
- Start a new chat for a new task instead of continuing an old, sprawling one; a fresh session gives every fact a full share of attention again.
- Restate the two or three facts a task depends on near the end of a long message rather than trusting they are still getting read clearly from ten screens back.
- Trim pasted documents and old tool output down to the excerpt that matters instead of leaving the full original text sitting in the conversation.
- Treat a chat transcript as disposable working space, not storage; anything that needs to survive past this session belongs somewhere that does not degrade with length.
A simple diagnostic: if an answer feels off, do not just rephrase the question in the same chat. Open a fresh conversation, paste in only the two or three facts that matter, and ask again. If the answer improves, length was the problem, not the question.
MemX's own guide to ChatGPT forgetting earlier messages covers the practical side of this in more depth: context window sizes by plan and tier, and the step-by-step settings to check before a conversation grows unwieldy. That page is the checklist; this post explains why the checklist works.
Where a Memory Layer Outside the Chat Helps
Attention dilution and positional degradation are both a property of what sits inside a single prompt that keeps growing. A fact typed on message four and a fact typed on message four hundred are both riding along in the same expanding block of text, subject to the same shrinking attention share and the same fuzzier position math the longer that block gets. A memory layer that lives outside the chat window sidesteps that failure mode differently: instead of keeping a fact parked inside the prompt, it stores the fact separately and pulls it back in through a fresh retrieval step only when it is needed. Retrieval happens the same way whether it's message one of a new chat or message four hundred of a long one, because the fact was never carried inside that block of text in the first place. MemX is built around this specific, narrow problem: keeping the facts that matter outside a conversation's own token count, private by architecture through per-user isolation, encryption at rest, and on-device processing where the platform supports it, so a fact's odds of being read correctly do not fall simply because a conversation happens to run long.
01Why does ChatGPT (or any AI chat) get dumber the longer you talk to it?
Two mechanisms compound. More tokens dilute attention, because every new message adds tokens that compete for a fixed attention budget, thinning the share any single fact receives. Separately, the positional encoding models use to track word order relies on angle-based math that gets fuzzier at long range, making it harder for the model to judge how far apart two pieces of information sit. Neither effect requires the conversation to be anywhere near the token limit, which is why it can feel like the model is getting dumber well before it runs out of room.
02What is context rot?
Context rot is the pattern where a large language model's accuracy on the same task falls as the surrounding input grows longer, even while the model is nowhere close to its maximum context window. Chroma documented it across 18 frontier models on tasks as simple as copying text back verbatim, ruling out the idea that longer inputs just happen to contain harder questions.
03Is context rot the same thing as running out of the context window?
No. Running out of the context window is a hard limit: the model literally cannot see a message once it scrolls past the input boundary. Context rot happens well before that limit, often at a small fraction of it, and shows up as gradually worse answers rather than a message the model can no longer read at all.
04Does starting a new chat fix it?
In most cases, yes, because a fresh conversation resets both mechanisms: fewer tokens means each one gets a larger share of attention, and every position is close to the start, where the model's positional math is most reliable. Restating the handful of facts the task depends on in that new chat works better than continuing to add messages to an old one.
05Does context rot happen in Claude and Gemini too, or only ChatGPT?
It happens across every major model family tested. Chroma's report covered GPT, Claude, Gemini, and Qwen models and found every one of them degraded as input length grew. The NoLiMa study separately found GPT-4o, Claude 3.5 Sonnet, Gemini 2.5 Flash, and Llama 4 Scout all lost significant accuracy on inference-based retrieval once input reached roughly 32,000 tokens. It is a property of how current transformer architectures process long input, not a flaw specific to one company's model.
