Why do some AI models get worse as you feed them longer inputs, even within their stated limits?
Why do some models like Llama degrade under long context?
This explores why language models (Llama is the example in the question) get worse as their input gets longer, even when the input still fits inside the advertised context window. The corpus has no study specific to Llama, so this answer covers the general mechanisms that apply to Llama and similar models.
This explores why models like Llama get worse as their input grows, even when it still fits within the advertised context window. One caveat first: the collection has no study of Llama's long-context behavior specifically. What it does have is a set of findings that, read together, explain the general pattern. The most surprising one is how early the decline starts. In one study, reasoning accuracy fell from 92% to 68% after only about 3,000 tokens of padding were added. That is a small fraction of what most context windows hold. The drop appeared across different tasks, did not track how good the model was at ordinary language modeling, and did not go away with chain-of-thought prompting Does reasoning ability actually degrade with longer inputs?. So the context window size tells you how much text a model can accept, not how much it can actually reason over.
Why does this happen? One view is that the problem isn't storage at all. Holding text in context is cheap. The expensive part is the computation needed to turn that text into something the model can use. When researchers gave models extra offline passes to absorb evicted context into their weights, performance kept improving with more passes, the same way extra thinking time helps on hard problems Is long-context bottleneck really about memory or compute?. A second mechanism compounds the first. In long tasks, the context fills with the model's own earlier output, mistakes included. Once errors are in the history, the model starts conditioning on them and makes more errors, and the decline is non-linear. Notably, bigger models don't fix this; only 'thinking' models that reason before answering reduce it Do models fail worse when their own errors fill the context?. In conversations, part of what looks like long-context decline is really the model losing track of what the user wants. Its training rewards answering early instead of asking clarifying questions Why do language models lose performance in longer conversations?.
Long context doesn't fail everywhere, though, and the pattern of where it holds up is informative. Long-context models can match retrieval-based (RAG) systems at finding semantically relevant passages, but they break down on structured queries that require joining information across tables Can long-context LLMs replace retrieval-augmented generation systems?. Feeding a strong reader larger chunks of a few thousand tokens can even beat precise retrieval of short snippets Can long-context models resolve retriever-reader imbalance?. Long context works well for skimming for relevance. It does poorly when the model has to work through many pieces of the input together.
The most interesting response in the collection is to stop pushing everything through attention. Recursive Language Models store a long prompt in a code environment and let the model query it piece by piece. They handle inputs about 100× beyond the context limit and do better than the base model even on short prompts Can models treat long prompts as external code environments?. ReadAgent copies how people read: it compresses a document into short 'gist' memories and goes back to the original only when it needs details. This extends usable context 3–20× Can LLMs read long documents like humans do?. Agent harnesses take the same idea further with layered memory: weights, context, a working scratchpad and disk Can external state caches let models solve harder problems?.
The finding you probably didn't know you wanted: how a model degrades changes with its capability. Weaker models visibly drop content from long documents. Frontier models instead corrupt content quietly while the document still looks intact Does model capability change how documents degrade?. A better model may not degrade less over long inputs; its failures may just be harder to spot.
Sources 10 notes
FLenQA shows reasoning accuracy drops from 92% to 68% at just 3000 tokens of padding, far below context window capacity. The degradation is task-agnostic, uncorrelated with language modeling performance, and persists even with chain-of-thought prompting.
Research shows the bottleneck is not memory capacity but the compute required to consolidate evicted context into fast weights during offline sleep phases. Performance improves with more consolidation passes, following a test-time scaling pattern on harder reasoning tasks.
Error accumulation in context causes non-linear performance degradation in long-horizon tasks. Model scaling does not fix this; only test-time compute through thinking models reduces the effect by preventing error-contaminated context from biasing reasoning.
LLMs degrade in multi-turn settings because RLHF training rewards premature answers over clarification-seeking, creating pragmatic mismatch with individual user behaviors. A Mediator-Assistant architecture that explicitly parses user intent before execution recovers lost performance without retraining.
The LOFT benchmark shows LCLMs match RAG on semantic retrieval without explicit training, but cannot execute relational queries requiring joins across structured tables. Context length alone cannot bridge this gap.
Show all 10 sources
LongRAG shows that 4K-token units and long-context readers outperform 100-word retrieval on standard benchmarks. The optimal RAG design shifts from precise retrieval to coarse ranking plus deep reading as context windows expanded.
Recursive Language Models store long prompts in a Python REPL and query them via code execution, avoiding attention degradation. RLMs outperform base models even on shorter prompts while handling inputs two orders of magnitude beyond context windows.
ReadAgent compresses documents into gist memories before knowing the task, then retrieves details only when needed, extending effective context 3–20× and outperforming retrieval baselines on long-document QA.
Prime Agent organizes persistent state in four levels (weights, context, persistent REPL plus subagents, disk-backed history) to let models read and write addressable state beyond their instruction stream. The approach isolates harness failures from model failures and reported gains on ARC-AGI-3, though specific components remain unablated.
DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Long-context LLMs Struggle with Long In-context Learning
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
- Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
- A Human-Inspired Reading Agent with Gist Memory of Very Long Contexts
- Recursive Language Models
- FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions