Why isn't a 'clean read' enough proof that an AI's shrunk memory kept what actually mattered?
Why does context compression need reward signals beyond explicit coherence metrics?
This explores why, when an AI system shrinks or summarizes its context to fit more in, checking that the compressed version 'reads coherently' isn't a good enough test, and what other signals tell you the compression actually kept what matters.
This explores why a fluent, coherent summary of a long context can still be a bad compression, and what kinds of feedback would catch that. The corpus has no single note on training context compressors with rewards. It does approach the question from several sides, and together they make a clear case: coherence measures how the compressed text reads, not whether it still does its job.
Start with what compression is supposed to keep. The Titans architecture decides what goes into long-term memory by how surprising each token is, not by how well it fits the story so far Can neural memory modules scale language models beyond attention limits?. That reverses the coherence instinct. The details that disrupt a smooth narrative are often the ones worth keeping, and a compressor rewarded for coherence will tend to smooth them away. A related line of work argues that the real long-context bottleneck isn't storage. It's the compute spent turning evicted context into internal state, and results improve with more consolidation passes Is long-context bottleneck really about memory or compute?. If compression is work that can be done well or badly, you need a way to grade that work, and a fluency check can't do it.
The most useful grade is whether the task still comes out the same. Atom of Thoughts throws away the reasoning history at each step and keeps only the current reduced problem, and its correctness test is that the answer stays the same, not that the trace stays readable Can reasoning systems forget history without losing coherence?. That points to one kind of signal beyond coherence: does the compressed state still produce the right answer? There's also a limit on how much compression can do at all. Transformers provably beat state-space models at copying and retrieval because a fixed-size state can't hold arbitrary detail Can state-space models match transformers at copying and retrieval?. Any compressor makes that tradeoff, and only task-level feedback shows where it went wrong.
The reward-design literature explains why one overall score, coherence included, falls short. Holistic reward models overfit to surface features, and splitting quality into checkable sub-criteria fixes much of that Can breaking down instructions into checklists improve AI reward signals?. Applied to compression, that means asking concrete questions: did the key entity survive? Did the constraint survive? Did the number survive? Numerical rewards also stall because they don't say *why* something failed. Written critiques break those plateaus Can natural language feedback overcome numerical reward plateaus?, and reward models that reason before scoring raise the ceiling further Can reward models benefit from reasoning before scoring?. The most directly transferable idea comes from LongTraceRL. It mines process rewards from what search agents *read but didn't cite*, which are the hardest distractors, and it only rewards reasoning that reached a correct answer, so the model can't earn reward with a convincing story Can search agent behavior yield reliable process rewards for reasoning?. That is roughly what a compression reward needs: a signal for what to drop as well as what to keep, tied to downstream correctness.
There's one more reason coherence can't be the target. Even when information *is* in the context, models often ignore it if it conflicts with what they learned in training Why do language models ignore information in their context?. A compressed context can be accurate and coherent and still change nothing about what the model does. Reaching the context window and being used are separate outcomes, and only rewards measured on the model's behavior can tell them apart.
Sources 9 notes
Titans architecture separates attention (short-term, quadratic) from neural memory (long-term, compressed), prioritizing surprising tokens for storage. The model outperforms standard Transformers and linear RNNs across tasks while scaling to 2M+ token contexts without quadratic penalties.
Research shows the bottleneck is not memory capacity but the compute required to consolidate evicted context into fast weights during offline sleep phases. Performance improves with more consolidation passes, following a test-time scaling pattern on harder reasoning tasks.
Atom of Thoughts decomposes problems into DAGs and contracts them iteratively, ensuring each state depends only on the current problem—not prior steps. This memoryless approach eliminates historical baggage that bloats reasoning while maintaining answer equivalence.
Two-layer transformers can copy exponentially long strings while state-space models are fundamentally limited by their fixed-size latent state. Empirically, transformers dramatically outperform SSMs at copying and context retrieval in both synthetic and pretrained settings.
RLCF and RaR methods decompose instruction quality into verifiable sub-criteria, improving performance on benchmarks like FollowBench and HealthBench. This decomposition principle reduces overfitting to superficial artifacts that plague holistic reward models.
Show all 9 sources
Critique-GRPO shows that models stuck on performance plateaus can generate correct solutions when given chain-of-thought critiques, revealing that numerical rewards lack critical information about why failures occur and how to improve.
Three independent teams (RRM, RM-R1, DeepSeek-GRM) discovered that adding chain-of-thought reasoning before reward scoring enables adaptive test-time compute scaling for evaluation. Reasoning-based approaches raise the capability ceiling of reward models beyond what outcome-based evaluation achieves.
LongTraceRL mines entity-level reasoning signals from what search agents read but don't cite—the hardest distractors—and applies rubric rewards only to correct answers, structurally blocking reward fabrication while capturing intermediate reasoning quality.
Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Reward Reasoning Model
- RM-R1: Reward Modeling as Reasoning
- Repeat After Me: Transformers are Better than State Space Models at Copying
- Titans: Learning to Memorize at Test Time
- Understanding and Mitigating Premature Confidence for Better LLM Reasoning
- Reasoning Language Models: A Blueprint
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs