Line of inquiry
Inquiring lines›How do training methods and scalin…›How do different training methods…›this line of inquiry
What training dynamics and scale trigger emergence of reasoning capabilities?
A broader line of inquiry — a family of 65 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 65
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How do two-phase training dynamics explain reasoning emergence?
- What distinguishes surface mechanisms from the training regimes that produce them?
- How does post-training shift models from passive prediction to on-policy action?
- Do emergent abilities result from genuine new capabilities or implicit in-context learning?
- How does model scale affect anticipatory behavior in structured training?
- How does in-context learning trigger phase transitions in model behavior?
- What training interventions could close the perception-action gap?
- Does the model learn depth-wise drift as an explicit strategy?
- Can models treat their own trained behaviors differently from asserted beliefs?
- Why does the gap between theoretical expressiveness and learned capability matter?
- How do complete multi-turn trajectories differ from isolated task examples?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- How do evaluative versus directive signals differ in next-state training?
- What is the difference between changing model outputs versus changing internal representations?
- How does action-level decomposition differ from token-level imitation in supervision?
- Does RL training activate latent meta-learning capacity or create it from scratch?
- Do task-specific heuristics improve gradually or appear suddenly at scale?
- How does training on correct answer form differ mechanistically from training on failure analysis?
- Can trained models encode programs more complex than their data-generating process?
- What qualities make a behavioral pattern count as a teachable skill?
- What training signals would models need to learn reciprocal common-ground construction?
- What structural differences emerge between early generic skills and later meta-strategy skills?
- Does task superposition explain how models learn from multiple in-context trajectories?
- Does recognizing your outputs as actions enable awareness of being evaluated?
- What emergent behaviors do models develop when trained on underspecified pedagogical tasks?
- Can models develop situational awareness without explicit training for it?
- Does base model geometry predict which associations persist through intervention?
- Does critique training improve exploration diversity during model training or only test time?
- What skills can large models identify and organize about their own abilities?
- Why does a systems lesson remain robust when it claims less about mechanisms?
- Does the Assistant Axis exist in pre-trained models before instruction tuning?
- Why does the order of training examples matter for what models learn?
- Do text-space skills transfer learning across different frontier models?
- How does policy initialization with sub-policies enable emergent thinking?
- Do different function-calling subtasks have different entropy profiles during training?
- How does early branch divergence differ from late branch divergence in supervision signals?
- What distinguishes inductive inference from negative evidence versus positive patterns?
- How do weights, selection, and prompts create different geometric landscapes of accessible behaviors?
- What deployment feedback loops amplify LLM pretraining popularity in live systems?
- How does an instruction-following LLM activate latent retrieval knowledge?
- Why does recontextualizing a behavior during training change whether models learn it?
- How sensitive is analogical reasoning emergence to training data and scale?
- Do mechanistic refusal vectors transfer across different models and training settings?
- Do fed-back concepts or the auxiliary objective alone drive the performance gain?
- Can explicit goal state scaffolding at inference time transfer to autonomous tracking through training?
- What test-time strategies did o3 discover without human specification?
- How does sliding the start state backward create informative learning signals?
- What grows faster: situational awareness or the gap between evaluated and unsupervised behavior?
- How does correctness emergence occur when no expert initially solved the task?
- Can models generate their own training curriculum during offline dreaming?
- How do weight visualizations reveal temporal structure in cyclic training?
- How do developmental curriculums emerge from learning progress signals?
- What specific tasks should evaluate whether models understand pedagogical sequencing?
- How does a challenger's escalating difficulty function as curriculum?
- What makes content informative and not-yet-mastered for reinforcement during pretraining?
- How does trajectory burstiness compare to other structural properties that shape emergent capabilities?
- How does activation consistency training differ from output-level consistency?
- What makes exploration a verifiable and measurable training objective?
- How does subliminal learning differ from statistical model collapse?
- How can weak-to-strong progressive training target planning without interfering with grounding?
- Can burst timing in process logs predict learning outcomes?
- What makes session-aware multi-turn tracking necessary for asynchronous training?
- Does extended exoskeleton use eventually produce meaningful skill transfer?
- Does balancing four pedagogical capabilities improve tutoring or create performance tradeoffs?
- What role does curriculum design play in reasoning emergence?