Line of inquiry
Inquiring lines›How should we train models for cap…›What systematic failures and vulne…›this line of inquiry
Why does finetuning cause catastrophic forgetting of model capabilities?
A broader line of inquiry — a family of 29 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 29
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does finetuning facts into weights overwrite existing model capabilities?
- How does in-weight memorization scale with model parameter count?
- Why does in-weight memorization fail compared to tool-based fact access?
- What makes factual memorization less efficient than tool-based retrieval?
- How does in-weights adaptation create spurious forgetting in models?
- How do layer-wise versus parameter-wise merging strategies affect information retention?
- Why does fine-tuning models for continuous reasoning cause catastrophic forgetting?
- How tight should a textual learning rate be before it prevents skill escape?
- What mechanism transfers explicit memories into parametric model weights?
- Does sparse parameter updating improve test-time training's computational cost?
- How do trained weights differ from a stored library or text?
- How do newly learned facts become accessible after gradient updates?
- Can structural perturbations harm model accuracy more than semantic ones?
- Does parameter isolation per task enable online updates without retraining?
- What causes overfitting when forcing new facts into model weights?
- Why does tool use decouple factual capacity from model parameter count?
- How should skill libraries coordinate with gradient-based weight optimization?
- How does weight sharing compound the advantages of deeper model designs?
- Where does skill extraction fail compared to genuine model adaptation?
- What makes representation interventions more efficient than weight perturbations for finetuning?
- Can we unlearn memorized text by finetuning only high-gradient weights?
- What makes knowledge editing different from simply finding where facts are stored?
- Can time-awareness live in model parameters instead of retrieval?
- How do byte-level representations enable better handling of typos than tokens?
- Does parameter composition work when adapter alignment is imperfect?
- Do different model sizes show different rates of optional field overfilling behavior?
- Why is editing specific facts so difficult in language models?
- Can dynamic variance weighting replace fixed objective combination weights?
- When should full-parameter post-training be used instead of LoRA adaptation?