Line of inquiry
Inquiring lines›How do training and design choices…›What determines whether training i…›this line of inquiry
What makes distillation transfer some model capabilities while suppressing others?
A broader line of inquiry — a family of 27 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 27
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- When does knowledge distillation produce student models superior to teachers?
- How can distillation preserve uncertainty expression instead of optimizing it away?
- Why does distillation transfer reasoning patterns with few examples?
- Does teacher scale matter for on-policy distillation success?
- What makes policy self-distillation more effective than external teacher distillation?
- Why does unsupervised self-distillation lose its advantage in thinking model mode?
- How does distilling only inconsistent rollouts compare to distilling all generations?
- Why does self-distillation suppress epistemic verbalization in student models?
- How does self-distillation differ from standard fine-tuning approaches?
- What causes on-policy distillation to become unstable at scale despite dense rewards?
- Why is offline knowledge distillation preferred when in-session signals matter?
- How do failure examples improve distillation compared to successful trajectories alone?
- What makes student-teacher distributional mismatch derail on-policy distillation?
- Does reasoning style transfer matter more than solution correctness in distillation?
- How does self-distillation degrade reasoning by suppressing uncertainty signals?
- Why does combining reasoning distillation with RLVR outperform either training stage alone?
- Can teachers trained under uncertainty constraints distill better generalizing students?
- Does distillation strip away uncertainty signals that reasoning actually needs?
- Can self-distillation reduce catastrophic forgetting in continual learning?
- How does prompt diversity compare to per-problem sampling depth in distillation?
- Why should we ignore bits where teacher and student already agree?
- Why does style transfer happen during knowledge distillation?
- Can ensemble predictions be distilled back into a single deployable model?
- What makes asymmetric distillation effective for converting pretrained diffusion models?
- Who serves as the teacher model in the routing-guided distillation process?
- How does upward distillation transfer knowledge from smaller to larger networks?
- Can experimental outcomes be reliably distilled into reusable insights?