Line of inquiry
Inquiring lines›How should we train models for cap…›What systematic failures and vulne…›this line of inquiry
What makes weaker teacher models effective for stronger student training?
A broader line of inquiry — a family of 31 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 31
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- When does knowledge distillation produce student models superior to teachers?
- Why do weaker teacher models sometimes produce better training signals than stronger ones?
- Can self-training drift be prevented by applying student compatibility filtering?
- How can distillation preserve uncertainty expression instead of optimizing it away?
- Can weak models supervise the alignment of stronger models effectively?
- How does information asymmetry between teacher and student create the learning signal?
- What makes student-teacher distributional mismatch derail on-policy distillation?
- Why do weaker models generate better training data than stronger models?
- Can teachers trained under uncertainty constraints distill better generalizing students?
- Why is offline knowledge distillation preferred when in-session signals matter?
- Why should we ignore bits where teacher and student already agree?
- What causes on-policy distillation to become unstable at scale despite dense rewards?
- Does gradient-based influence estimation identify which alignment examples actually matter most?
- What makes asymmetric distillation effective for converting pretrained diffusion models?
- Can we cheaply estimate which samples are currently most informative?
- Can gradient-based influence scores beat difficulty metrics for identifying valuable training data?
- Why does teacher-student proximity matter more than absolute teacher strength?
- How does student capacity limit what it can learn from teachers?
- Can gradient-based influence estimation make test-time training more efficient?
- What alignment procedures cause different models to share the same output distribution?
- What filtering criteria best identify student-compatible refinements from teacher models?
- How does training data distribution create asymmetric competence across relation types?
- Why does style transfer happen during knowledge distillation?
- Why does information asymmetry between teacher and student enable effective feedback learning?
- How does subliminal learning differ from statistical model collapse?
- Why does training single-step consistency models prove so difficult compared to diffusion?
- How does upward distillation transfer knowledge from smaller to larger networks?
- How does activation consistency training differ from output-level consistency?
- How can weak-to-strong progressive training target planning without interfering with grounding?
- How does Easy Consistency Tuning accelerate consistency model training from diffusion checkpoints?
- Can signal quality regulations help smaller teachers outperform larger ones?