Line of inquiry
Inquiring lines›How should we train models for cap…›What systematic failures and vulne…›this line of inquiry
How does memorization interact with learning and generalization?
A broader line of inquiry — a family of 19 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 19
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why is extracting training data insufficient proof that models memorize?
- Can document repetition accidentally memorize sensitive information instead of learning?
- How much RLVR improvement comes from benchmark data memorization?
- Why does training data not function as a searchable corpus?
- How much does memorization capacity limit a model's ability to learn new information?
- How much training data teaches retrieval models to follow instructions?
- Why does test accuracy improve after training accuracy reaches 100 percent?
- What makes memorized paragraphs harder to corrupt than generic text?
- How much improvement comes from caching versus actual capability gain?
- How do out-of-distribution tests reveal that optimization learning is memorization?
- What is the theoretical capacity limit before memorization saturates?
- Can curated demonstrations compensate for smaller or simpler training environments?
- Why do older datasets show higher LLM performance than newer ones?
- How can post-training research become reproducible without releasing full interfaces?
- Can experimental outcomes be reliably distilled into reusable insights?
- Why do energy-based models generalize better on out-of-distribution data than standard transformers?
- Can vector store deletion truly prevent information recovery?
- How do training data cutoffs produce false claims that stay consistent?
- Why does curriculum order matter when information theory says data order is irrelevant?