Line of inquiry
Inquiring lines›What enables robust retrieval and…›How can memory and attention syste…›this line of inquiry
Why does adding new knowledge through fine-tuning degrade existing capabilities?
A broader line of inquiry — a family of 61 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 61
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does moving memory outside model weights avoid the limitations of in-weight retention?
- How does in-weight memorization scale with model parameter count?
- Can in-weight memorization scale beyond model parameter count limits?
- Why does fine-tuning models for continuous reasoning cause catastrophic forgetting?
- Can models recover knowledge with completely unrelated retraining tasks?
- Why does fine-tuning for continuous space cause catastrophic forgetting?
- How does in-weights adaptation create spurious forgetting in models?
- Does finetuning facts into weights overwrite existing model capabilities?
- Why does in-weight memorization fail compared to tool-based fact access?
- What makes factual memorization less efficient than tool-based retrieval?
- Can data pruning strategies exploit the finite nature of memorization capacity?
- What mechanism transfers explicit memories into parametric model weights?
- How do newly learned facts become accessible after gradient updates?
- Why does specializing to one task make future task learning harder?
- Why do pretrained model priors reduce the usefulness of retrieved experience?
- Is forgetting in language models reversible or permanent knowledge loss?
- What capacity limits does the memory model face as corpus grows?
- Why do large language models outperform fine-tuned models once repeated items are removed?
- Can externalized memory and skills replace model scaling?
- Why does semantic deduplication reduce memorization in fine-tuned models?
- What causes catastrophic forgetting during domain knowledge embedding?
- Why is extracting training data insufficient proof that models memorize?
- Do models with unfilled memorization capacity appear to generalize falsely?
- What makes some contexts learnable as rules versus requiring model retraining?
- Can document repetition accidentally memorize sensitive information instead of learning?
- How do layer-wise versus parameter-wise merging strategies affect information retention?
- How does parametric knowledge sabotage context-grounded question answering?
- Can models internalize retrieved context as static parametric knowledge?
- How do trained weights differ from a stored library or text?
- Can native memory procedures acquired through training handle stale or incorrect cached information?
- When does training a memory model beat RAG or fine-tuning?
- How can a forgetting policy preserve rare knowledge while preventing over-generalization?
- Can AI models retain knowledge across changing environments without catastrophic forgetting?
- Can we unlearn memorized text by finetuning only high-gradient weights?
- When does knowledge activation fail across different model architectures?
- How does distributional shift toward rare inputs change memorization reliance?
- What determines whether accumulated state generalizes spuriously across continual learning domains?
- How does memorization capacity saturation trigger the grokking transition?
- Can zero-weight drift through external memory replace parameter plasticity entirely?
- How much does memorization capacity limit a model's ability to learn new information?
- Why does fine-tuning fail to remove temporal contamination from pretraining?
- How do models develop dense representations for familiar training data?
- What causes overfitting when forcing new facts into model weights?
- Why does grokking reveal the shift from memorization to genuine understanding?
- Do sample-level similarities between pretraining and downstream tasks explain the frequency effect?
- Can models consolidate context into weights during idle offline phases?
- Does grokking in modular arithmetic follow the same three-phase learning trajectory?
- How do training-time and inference-time knowledge injection techniques compare?
- What is the theoretical capacity limit before memorization saturates?
- How can memory shift from a passive datastore to an actively trained component?
- What makes memorized paragraphs harder to corrupt than generic text?
- How does KL regularization prevent both forgetting and adaptation loss?
- Why does test accuracy improve after training accuracy reaches 100 percent?
- What explains the contextual variability of knowledge in transformers?
- What makes knowledge editing different from simply finding where facts are stored?
- How do early layers preserve unbiased information while late layers conform?
- How do out-of-distribution tests reveal that optimization learning is memorization?
- Can a memory module be swapped between different base models?
- How would you redesign context integration to prevent prior associations from dominating?
- How do the three grokking phases connect to memorization capacity limits?
- Why does keyword priming require only three training exposures to establish?