Line of inquiry
Inquiring lines›How do training choices shape mode…›How do systems prioritize structur…›this line of inquiry
What persistent memory architectures support learning across sessions?
A broader line of inquiry — a family of 79 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 79
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do long-term memory modules outperform consolidation into fast weights?
- Can precomputed inferences be stored in memory modules between model interactions?
- What persistent memory architectures best support storing precomputed inferences across sessions?
- Does including full context always degrade memory retrieval quality in practice?
- Why do accumulated memory systems sometimes hurt continual learning?
- Can data pruning strategies exploit the finite nature of memorization capacity?
- How do adaptive memory modules compare to feedback-based working memory for long context?
- How do retention gates regularize forgetting across different sequence model architectures?
- How should memory consolidation timing differ across multiple timescales?
- How do complementary learning systems explain the need for fast and slow consolidation?
- How should memory systems split between short-term and long-term storage?
- Why does attending to own latents work better than bolted-on external memory stores?
- How does context budget create tradeoffs between memory and skills?
- Why does fine-tuning for continuous space cause catastrophic forgetting?
- What makes factual memorization less efficient than tool-based retrieval?
- Can continuum memory systems prevent catastrophic forgetting in neural networks?
- Do retrieval-augmented memory systems actually solve the compartmentalization problem?
- Why does specializing to one task make future task learning harder?
- Why does semantic deduplication reduce memorization in fine-tuned models?
- How do fixed recurrent states trade off copying accuracy for filtering ability?
- Can adaptive memory modules combine long-term filtering with short-term attention benefits?
- How do memorization and attention map onto different memory systems?
- What gets lost when we describe memory as retrieval?
- How does in-weight memorization scale with model parameter count?
- Do models with unfilled memorization capacity appear to generalize falsely?
- Why does in-weight memorization fail compared to tool-based fact access?
- Can neural modules memorize surprising tokens as adaptive long-term memory?
- When does training a memory model beat RAG or fine-tuning?
- Why does persistent memory alone fail to create genuine position-holding in models?
- Why does recency-based recall outperform semantic similarity for episodic memory?
- How can a forgetting policy preserve rare knowledge while preventing over-generalization?
- Can memory primitives become first-class design objects like computation sparsity?
- Can episodic and semantic memory improve long-horizon task reasoning?
- Why is extracting training data insufficient proof that models memorize?
- Can document repetition accidentally memorize sensitive information instead of learning?
- How do cortical columns implement local inference over memory cycles?
- Can memory-based adaptation and gradient fine-tuning operate on complementary timescales?
- How does distributional shift toward rare inputs change memorization reliance?
- What makes a learned consolidation rule lossy and where does contamination enter?
- How does memorization capacity saturation trigger the grokking transition?
- Why do language models ignore condensed memory even when it is the only memory?
- Can models consolidate context into weights during idle offline phases?
- What determines whether accumulated state generalizes spuriously across continual learning domains?
- Why should consolidation be scheduled offline rather than during forward passes?
- Why does storing past judgments in memory make current evaluations worse?
- Can we unlearn memorized text by finetuning only high-gradient weights?
- How can memory shift from a passive datastore to an actively trained component?
- Can offline recurrent passes replicate sleep-based memory consolidation in AI?
- How much does memorization capacity limit a model's ability to learn new information?
- Why does grokking reveal the shift from memorization to genuine understanding?
- Why is consolidation quality the binding constraint in neural memory systems?
- How do retrieved memories differ from decision-context passages for prediction?
- How does consolidation schedule order affect final memory quality?
- What makes structured memory schemas more stable than freeform text summaries?
- Why does a replay mechanism prevent reasoner skills from over-specializing?
- Does grokking in modular arithmetic follow the same three-phase learning trajectory?
- Why are rare tokens the hooks for verbatim model memorization?
- How do the six memory components combine across explicit and implicit paths?
- How do memory hierarchies and compression reduce context management demands?
- What is the theoretical capacity limit before memorization saturates?
- Can test-time scaling compound through memory consolidation into a new scaling law?
- How does dual-rate learning separate episodic and procedural memory in neural networks?
- Can episodic raw memory outperform consolidated summaries in practice?
- Why do cognitive metaphors change based on available technology?
- What triggers control processes to act on stored preference knowledge?
- How do the three grokking phases connect to memorization capacity limits?
- What makes a memory reachable in the right context?
- What computational costs does closed-loop memory refinement introduce?
- How does KL regularization prevent both forgetting and adaptation loss?
- How do out-of-distribution tests reveal that optimization learning is memorization?
- How should memory systems handle deletion as a structural property?
- Why do hybrid memory systems outperform single-tier AI architectures?
- What makes memory trajectories topologically stable under persistent reuse?
- What makes knowledge seeding equivalent to hippocampal replay in the brain?
- Why does connectivity between memory modules matter more than storage capacity?
- How does continuous implicit memory formation differ from explicit memory encoding?
- How should we measure operational cost of memory systems in production?
- How does co-activation shape which memories become linked together?
- How does the hippocampus bind disparate elements without storing everything itself?