Line of inquiry
Inquiring lines›What drives capability improvement…›How do training signals and method…›this line of inquiry
How do training data quality and composition affect downstream model performance?
A broader line of inquiry — a family of 86 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 86
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does teacher-style refinement of training data transfer equally to all student model distributions?
- At what point does output quality outweigh diversity value in synthetic data tasks?
- How do quality and diversity in synthetic data interact with accumulation schedules?
- Can selecting the right data subset outperform training on everything?
- How does the ratio of synthetic to real training data affect model collapse?
- How do quality, diversity, and complexity create different effects on downstream model performance?
- Why does diversity of training cases matter more than raw dataset size?
- Why does mixed instruction data sometimes hurt specific model capabilities?
- How does diversity loss in synthetic data mirror tail distribution disappearance?
- Does selecting examples from multiple complexity levels outperform selecting only high-quality examples?
- What conditions make training diversity better than individual expert quality?
- What makes training data quality more important than quantity for reasoning?
- Why does diversity in training data enable denoising rather than reinforce shared biases?
- Why do unified models still inherit data-distribution biases from training?
- Why do weaker models generate better training data than stronger models?
- What creates the irreducible trade-off between quality and diversity in training data?
- Why do proprietary models improve with training while open-source models decline?
- Why do easy training examples contribute less to model generalization than hard ones?
- Can scaling up contradictory training data overcome unpredictable override effects?
- Why does synthetic-only data degrade performance even when matched in size?
- Can decoding-time prompting strategies fully replace diversity-focused training methods?
- How does the Learning Law explain why all examples should contribute equally?
- Can synthetic data diversity preserve the transcendence effect or does it collapse?
- Can training data organization by capability outperform source or task-based mixing?
- How do complexity and diversity affect model performance differently?
- Why does the same training data produce different gains across models?
- What training data contamination rates threaten model safety most practically?
- Can synthetic data preserve the diversity needed for transcendence to work?
- When does knowledge distillation produce student models superior to teachers?
- How does graph-based tool sampling differ from random sampling in diversity?
- Does importance sampling actually recover capabilities lost to hard sample training?
- How much does training data composition shape security model performance?
- Can we cheaply estimate which samples are currently most informative?
- How does training data distribution determine what models can learn?
- Can data pruning and equal contribution be reconciled in optimal learning?
- Why do rare complex structures in training data harm LLM generalization?
- How does prompt diversity compare to per-problem sampling depth in distillation?
- Is model collapse a property of data replacement or synthetic data itself?
- How does training frequency distribution shape what models reliably retrieve?
- Does verbalized sampling preserve factual accuracy and safety during diversity gains?
- How much does diversity training cost in single-shot pass@1 performance?
- Why is evaluating synthetic data quality so ambiguous and context-dependent?
- Does debiasing training data actually solve the bias problem in machine learning?
- Can gradient-based influence estimation make test-time training more efficient?
- Why does separating global coverage from local variation improve synthetic data generation?
- What distinguishes instance seeds from full input-output exemplar requirements?
- Can gradient-based influence scores beat difficulty metrics for identifying valuable training data?
- Does teacher scale matter for on-policy distillation success?
- How do you verify whether your context distribution satisfies covariate diversity?
- What determines whether additional training data still improves model performance?
- Why do models fail under distribution shift if accuracy metrics stay high?
- Why is offline knowledge distillation preferred when in-session signals matter?
- Can expert vectors learned offline transfer across multiple model architectures?
- Can temperature-adaptive sampling reduce the sharpening tax during training?
- Can curated demonstrations compensate for smaller or simpler training environments?
- Can synthetic data generation work without seed examples?
- Can teachers trained under uncertainty constraints distill better generalizing students?
- How much does preference data freshness matter compared to data source in DPO?
- What makes seed data a bottleneck in synthetic generation pipelines?
- How does training data distribution create asymmetric competence across relation types?
- Do identical task structures mean repeated instances or new synthetic samples with same design?
- Can balancing training data by source eliminate agent source preference bias?
- What makes student-teacher distributional mismatch derail on-policy distillation?
- How does off-policy data reuse inside trust regions affect convergence guarantees?
- Does unpredictable generalization from SDF become predictable at different training document scales?
- Why do certain tokens at certain difficulties drive most of RLVR's learning signal?
- How much performance is lost when converting pretrained checkpoints versus training from scratch?
- Why does positive reinforcement degrade diversity at higher k values?
- How do cyclic learning rates anti-correlate with weight decay to create diversity?
- Can complexity, diversity, and fidelity scale together in synthetic environments?
- Why do energy-based models generalize better on out-of-distribution data than standard transformers?
- Can training data analysis predict which samples will cause unintended personality changes?
- What makes asymmetric distillation effective for converting pretrained diffusion models?
- Why should we ignore bits where teacher and student already agree?
- Why do moderately represented cultures show more flattening than data-poor cultures?
- Can seedless generation maintain explainability while scaling control?
- How does business logic specification replace annotated training datasets?
- Can information-gain principles improve how we choose what to label?
- How should training distribution distance be defined when the policy evolves?
- Why does verification sit on a different scaling axis than pre-training?
- When should full-parameter post-training be used instead of LoRA adaptation?
- What baseline evidence distinguishes amplification from unchanged failure rates?
- How do training data cutoffs produce false claims that stay consistent?
- Why does curriculum order matter when information theory says data order is irrelevant?
- How does modified PPO handle samples from much older model versions?
- How much does sliding-window augmentation improve single-session modeling?