Line of inquiry
Inquiring lines›How do language models construct a…›How does AI persuasion undermine h…›this line of inquiry
Why do continual learning scenarios trigger catastrophic forgetting and interference?
A broader line of inquiry — a family of 53 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 53
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do long-term memory modules outperform consolidation into fast weights?
- Why does specializing to one task make future task learning harder?
- Does latent density emerge during pretraining from training data familiarity?
- How do complementary learning systems explain the need for fast and slow consolidation?
- Why does fine-tuning for continuous space cause catastrophic forgetting?
- Why do large language models outperform fine-tuned models once repeated items are removed?
- What determines whether accumulated state generalizes spuriously across continual learning domains?
- Can data pruning strategies exploit the finite nature of memorization capacity?
- How does distributional shift toward rare inputs change memorization reliance?
- Do sample-level similarities between pretraining and downstream tasks explain the frequency effect?
- What inductive bias would force models to learn Newtonian mechanics instead of shortcuts?
- How does representational density emerge from training data familiarity?
- Can zero-weight drift through external memory replace parameter plasticity entirely?
- Do models with unfilled memorization capacity appear to generalize falsely?
- How does the Learning Law explain why all examples should contribute equally?
- How can a forgetting policy preserve rare knowledge while preventing over-generalization?
- Does representational density emerge from training data exposure during pretraining?
- Why does semantic deduplication reduce memorization in fine-tuned models?
- Can training order and structure shape what networks retain and learn?
- Does the model learn depth-wise drift as an explicit strategy?
- Can memory-based adaptation and gradient fine-tuning operate on complementary timescales?
- Is forgetting in language models reversible or permanent knowledge loss?
- Can AI models retain knowledge across changing environments without catastrophic forgetting?
- Does grokking in modular arithmetic follow the same three-phase learning trajectory?
- How do models develop dense representations for familiar training data?
- How does memorization capacity saturation trigger the grokking transition?
- How does training frequency distribution shape what models reliably retrieve?
- How do learning dynamics on one example shift predictions on other responses?
- Why should scaling laws be understood as properties of data distribution rather than training in general?
- How do overparameterization and data size shift what attractors represent?
- How does dynamic recurrence during training improve depth extrapolation?
- Does environment stochasticity force models to generalize better across trajectory variations?
- Can pretraining-frequency signals alone prevent RAG systems from confabulating about common knowledge?
- How does KL regularization prevent both forgetting and adaptation loss?
- Can the joint-training principle extend beyond memorization and generalization pairs?
- Do KANs maintain their advantages in deep architectures and large-scale training?
- Can latent recurrence overcome the trainability costs of depth?
- What distinguishes data that generalizes broadly from task-specific memorization?
- Should loop count be fixed at training time or selected at test time?
- What happens to representational structure during model pretraining phases?
- How do cyclic learning rates anti-correlate with weight decay to create diversity?
- Why do optimal learning dynamics improve scaling law coefficients specifically?
- Can gradient approximation at equilibrium replace backpropagation through time in practice?
- Can data pruning and equal contribution be reconciled in optimal learning?
- How does dual-rate learning separate episodic and procedural memory in neural networks?
- Does weight decay directly cause contractive behavior near training examples?
- What non-parametric methods could replace latent factors for inductive learning?
- Can self-distillation reduce catastrophic forgetting in continual learning?
- Why does curriculum learning with tight budgets beat fixed-budget approaches?
- What makes data augmentation an implicit form of contraction learning?
- Can autoencoders act as associative memory systems like Hopfield networks?
- How do the three grokking phases connect to memorization capacity limits?
- Why does recomputing weights cost less than moving them on phones?