Line of inquiry
Inquiring lines›How do training signals reliably a…›What training signals and data cur…›this line of inquiry
How do training data composition and selection affect model capabilities?
A broader line of inquiry — a family of 53 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 53
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How much does training composition affect syntactic versus reasoning performance?
- Can training on diverse related tasks be more efficient than task-specific training?
- How much task-similar finetuning data does test-time training actually need?
- How do task frequency and complexity interact with model capacity during training?
- Why does exploration quality matter more than learner network depth?
- Can selecting the right data subset outperform training on everything?
- How much of the combinatorial task space must training data cover?
- How do training data distributions constrain what language models can accurately know?
- Why does mixed instruction data sometimes hurt specific model capabilities?
- Can scaling data alone solve performance gaps on long-tail concepts?
- What makes training data quality more important than quantity for reasoning?
- Does knowledge structure matter more than knowledge volume for model training?
- How tight should a textual learning rate be before it prevents skill escape?
- Can training order and structure shape what networks retain and learn?
- How does training distribution shape what language models understand best?
- Can a single model trained on two tasks predict untrained decision tasks?
- Why should scaling laws be understood as properties of data distribution rather than training in general?
- What distinguishes data that generalizes broadly from task-specific memorization?
- Does training on granular tasks beat training on the full function calling problem?
- Can intentional data-mixture design replace model scaling for rare task learning?
- How should training data be constructed to preserve teacher-student information gaps?
- Why does semantic similarity retrieval enable skill transfer to novel situations?
- Why do structure-targeted training negatives fail to fix the underlying problem?
- Why do rare complex structures in training data harm LLM generalization?
- How do task difficulty and skill type interact in model performance?
- Do sample-level similarities between pretraining and downstream tasks explain the frequency effect?
- Why does training data not function as a searchable corpus?
- How does training frequency distribution shape what models reliably retrieve?
- How should skill libraries coordinate with gradient-based weight optimization?
- Why do non-experts default to familiar chart types despite domain complexity?
- How much training data teaches retrieval models to follow instructions?
- Does prompt performance vary by how well training data covers the domain?
- What makes structured stochasticity more effective than unstructured randomness in reasoning?
- Does environment stochasticity force models to generalize better across trajectory variations?
- Do identical task structures mean repeated instances or new synthetic samples with same design?
- What happens when a single loss function conflates representation learning with decision-making?
- Why do image captions create different friction than pure video data?
- What makes a task at the edge of competence optimal for RL?
- Why is offline knowledge distillation preferred when in-session signals matter?
- Why does combining natural language with numerical scores improve prediction accuracy?
- How does training data distribution create asymmetric competence across relation types?
- How does cognitive fit theory explain why different tasks need different knowledge structures?
- How much does multi-token prediction help in protein design specifically?
- Why should we ignore bits where teacher and student already agree?
- Why do text-to-image models fail at composing multiple concepts together?
- How does business logic specification replace annotated training datasets?
- Why does masking the penultimate token outperform random token masking?
- Why do sigmoid conflict curves look the same across different language models?
- What makes certain bond distributions more learnable than others?
- How many skills in a library create too many combinations to scan?
- Why does homework adherence remain low despite advances in language model capability?
- Why does curriculum order matter when information theory says data order is irrelevant?
- Do substitute networks converge differently than complement networks?