INQUIRING LINE

Should an AI compress knowledge with a bolted-on module, or is 'compression' just what training already does by itself?

Can a separate frozen module manage compression better than joint optimization?

This explores whether compressing knowledge or context works better when you hand the job to a separate, dedicated module sitting beside a frozen model, rather than training everything together so the model learns to compress for itself.


This explores whether a separate, dedicated module beside a frozen model compresses knowledge or context better than training everything together. One thing first: the corpus has no paper that runs the two approaches against each other. What it has is a set of separate-module systems that work well, plus a deeper line of work saying that training a language model already is compression. Read together, they suggest the answer depends on what you're compressing and who will be reading the result.

The strongest case for a separate module involves knowledge. Memory Decoder takes what a retrieval system would have looked up and squeezes it into a small transformer. That transformer plugs into any LLM by blending its predictions with the LLM's own. It keeps rare facts and cuts perplexity by about 6 points, with no search step when the model runs Can retrieval knowledge compress into a tiny parametric model?. MeMo makes a similar move. A dedicated memory model holds the new knowledge, so query cost doesn't grow with the size of the document collection, and the main model can stay frozen, even a closed proprietary one. The costs are training up front and a cap on how much the memory model can hold Can a separate memory model inject knowledge without touching the LLM?. A related result comes from finetuning. Training small edits to a frozen model's internal activations, instead of changing its weights, used 10 to 50 times fewer parameters than LoRA Can editing hidden representations beat weight updates for finetuning?. So a frozen core with a small trained part attached can be very efficient.

For compressing a conversation or task history, the most interesting result is that the right amount of compression depends on the model. AdaCoM trains an outside manager with reinforcement learning to decide what to cut and what to keep for a frozen agent. Strong agents did best when more detail was kept. Weak agents did best when the context was cut hard Can an external manager handle context for frozen agents?. That works partly in favor of a separate module, since you can fit it to whichever model it serves without retraining that model. It also gives the reason joint training appeals: compression only helps if it suits the model that reads the result. Another line of work says the hard part of long context is the compute needed to fold old context into the model's own weights, not storage space. Results improve with more of these consolidation passes Is long-context bottleneck really about memory or compute?. That is a case for the model doing the compression itself, if you can pay for the compute.

Here is the part you might not expect. Several papers argue that a trained language model already is a compressor. Models trained only on text compressed images and audio better than PNG and FLAC by adapting through their context window Can text-trained models compress images better than specialized tools?. Another paper derives the best possible training process from a lossless compression objective Does optimal language model learning maximize data compression?. If that's right, a separate module never replaces the model's own compression. It adds a second layer of compression on top of the first. That's why the separate-module approach does best where the frozen model's training can't reach: new facts (MeMo, Memory Decoder) and fitting compression to a particular agent (AdaCoM). It's also why it hits capacity limits. Harness work points the same way: improving the code and tools around a frozen model, rather than the model itself, produced large benchmark gains Can frozen models improve by evolving their harnesses?.


Sources 8 notes

Can retrieval knowledge compress into a tiny parametric model?

Memory Decoder successfully compresses kNN-LM retrieval distributions into a small transformer that plugs into any LLM via output interpolation. It preserves long-tail factual knowledge while maintaining semantic coherence, reducing perplexity by 6.17 points across domains.

Can a separate memory model inject knowledge without touching the LLM?

MeMo trains a dedicated memory model to encode new knowledge, eliminating inference-time search costs that scale with corpus size. It avoids fine-tuning risks and works with frozen proprietary models, but trades this for up-front training cost and capacity limits.

Can editing hidden representations beat weight updates for finetuning?

ReFT learns task-specific interventions on frozen model representations rather than updating weights, with LoReFT (low-rank linear subspace variant) dramatically outperforming LoRA across reasoning, instruction-following, and NLU benchmarks while using far fewer parameters.

Can an external manager handle context for frozen agents?

AdaCoM trains an external RL-based manager to prune and preserve context for frozen agents. The key finding: stronger agents benefit from high-fidelity preservation, while weaker agents need aggressive compression—optimal context management is agent-specific, not task-universal.

Is long-context bottleneck really about memory or compute?

Research shows the bottleneck is not memory capacity but the compute required to consolidate evicted context into fast weights during offline sleep phases. Performance improves with more consolidation passes, following a test-time scaling pattern on harder reasoning tasks.

Show all 8 sources
Can text-trained models compress images better than specialized tools?

Chinchilla models trained exclusively on text achieve better compression rates on images and audio than FLAC and PNG by using their context window to adapt as task-specific compressors. This demonstrates that generalization operates through compression, not specialization.

Does optimal language model learning maximize data compression?

Research shows that optimal LM training can be derived from a lossless compression objective, yielding a Learning Law where all examples contribute equally in the optimal process. This approach improves scaling law coefficients, not just constants.

Can frozen models improve by evolving their harnesses?

DarwinX achieves average 17-point gains across benchmarks by evolving harness variants (prompts, tools, skills, control flow) under a preserve-and-extend contract while keeping the model frozen. Key evidence includes Terminal-Bench 2.1 rising to 84.7% and WebArena-Infinity reaching 93.0% audit-clean pass@1.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.