Line of inquiry
Inquiring lines›How do training signals reliably a…›What training signals and data cur…›this line of inquiry
Do pretraining and finetuning change model capabilities or only output behavior?
A broader line of inquiry — a family of 45 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 45
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How do finetuning and pretraining improvements differ in their effects on model capabilities?
- Does fine-tuning actually change model capabilities or only output distribution?
- How does behavioral fine-tuning differ from factual knowledge encoding in models?
- What is the difference between changing model outputs versus changing internal representations?
- How does model scale affect anticipatory behavior in structured training?
- Why does the gap between theoretical expressiveness and learned capability matter?
- Does pretraining data size matter less than base model scale for finetuning?
- What happens to base model capabilities when you apply finetuning?
- Why does fine-tuning change how models process retrieved context?
- What capabilities actually require massive scale versus specialized training regimes?
- How does the functional separation of knowledge and reasoning affect adaptation methods?
- Why does the order of training examples matter for what models learn?
- How much performance is lost when converting pretrained checkpoints versus training from scratch?
- Why does recontextualizing a behavior during training change whether models learn it?
- What happens to representational structure during model pretraining phases?
- How much does pretraining contribute to ToM performance versus task-specific training?
- Why does training order matter across different domain types?
- Which finetuning method works best across different task and data regimes?
- How does pretrained knowledge constrain what adaptation strategies can achieve?
- What happens to model capability as weight sparsity increases during training?
- What trade-offs emerge between training objectives and model reliability?
- Does the Assistant Axis exist in pre-trained models before instruction tuning?
- How do early layers preserve unbiased information while late layers conform?
- How should guidance levels adapt as the model's capability boundary shifts?
- How much does workflow architecture matter versus raw model capability?
- Do different domains require different types of model investment?
- How much does pretraining quality affect the modularity of fine-tuned models?
- How do training objectives shape what a world model actually learns?
- What counts as full capability recovery versus partial restoration?
- How much does workflow architecture matter compared to raw model capability in forecasting?
- How does model capability relate to personality conditioning flexibility?
- What makes content informative and not-yet-mastered for reinforcement during pretraining?
- What distinguishes new associations from existing ones at the computational level?
- How does evaluation format change what we measure about model reasoning?
- How can weak-to-strong progressive training target planning without interfering with grounding?
- Why does verification sit on a different scaling axis than pre-training?
- Why does teacher-student proximity matter more than absolute teacher strength?
- How should training distribution distance be defined when the policy evolves?
- What role does KL penalty strength play in format selection?
- What separates bootstrapping gains from sustained self-improvement gains?
- How does Easy Consistency Tuning accelerate consistency model training from diffusion checkpoints?
- How does prior coding experience change the way students use vibe coding tools?
- How much does sliding-window augmentation improve single-session modeling?
- What makes two timescales better than one for minimizing weight movement?
- How do fully crossed experimental factors differ from partially varied scenario conditions?