Line of inquiry
Inquiring lines›How do training methods and scalin…›What enables reasoning capability…›this line of inquiry
Is reasoning capability latent in base models or created by post-training?
A broader line of inquiry — a family of 74 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 74
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do base models contain latent reasoning that minimal training can unlock?
- What latent reasoning capability do base models already possess before training?
- Can minimal training signals unlock latent reasoning capability in base models?
- Can models possess latent reasoning capability that training signals fail to unlock?
- What makes reasoning capability a pre-training rather than post-training phenomenon?
- Does the base model already contain latent reasoning capability?
- Can minimal training signals unlock reasoning already latent in pretrained representations?
- Does latent reasoning capability exist in base models before any training?
- Do base models truly possess latent reasoning capability?
- Can models reason at inference without specialized internal training?
- What mechanisms activate latent reasoning capabilities already present in base models?
- Can pretraining signals unlock latent reasoning that post-training merely activates?
- Can RL training teach models when to activate reasoning versus when to skip it?
- Can smaller amounts of diverse reasoning demonstrations replace exhaustive factual training data?
- Does token-level reasoning during pretraining improve general reasoning without task-specific supervision?
- Does targeting the edge of competence during RL pretraining unlock true reasoning gains?
- Can training improve reasoning coherence without improving actual correctness?
- What distinguishes reasoning activation mechanisms across different training methods?
- What is the distinction between teaching reasoning how versus when to activate?
- How do single training examples activate reasoning capabilities in language models?
- Why do single examples trigger large reasoning improvements in models?
- Can distillation from stronger models create genuinely new reasoning abilities?
- How do reasoning training methods sacrifice some thinking skills while improving others?
- What makes some reasoning strategies genuinely novel versus latent?
- How much training data is truly necessary to unlock latent model reasoning?
- How can one training example improve reasoning across thousands of unseen problems?
- Can models learn to select exemplars based on reasoning skills rather than complexity?
- What other triggers can activate the latent reasoning capability?
- How does critique fine-tuning on one problem unlock broader reasoning?
- Does looped pretraining build reasoning more efficiently than supervised fine-tuning?
- How does backward reasoning during training improve forward reasoning capability?
- Can activation-space steering vectors replicate thinking model performance without retraining?
- Can targeted activation steering surface latent reasoning in base models?
- Why does reasoning training improve math but hurt knowledge tasks?
- Why does reasoning backward enable better forward reasoning performance?
- Why does extended reasoning training improve exploration without adding new capabilities?
- What makes reasoning-specific post-training different from standard parameter scaling?
- Why does imitation learning create a ceiling for reasoning capability?
- Can we predict when a model will develop thinking behaviors?
- How do timing and search internalization interact during reasoning post-training?
- Why do knowledge and reasoning train in different network layers?
- Does RL training actually restore the critical thinking that reasoning models lose?
- Can structured questioning prompts improve reasoning beyond standard conversational training?
- How does RPT compare to learning when versus how to deploy reasoning?
- How does a single training example trigger phase transitions in reasoning output?
- What makes thought identifiability provable without auxiliary training data?
- Why does pre-training provide the raw material for emergent thinking?
- Can thought quality alone be trusted to guide model training?
- What makes the verifier the load-bearing component of reasoning training?
- What makes token-level reasoning during pretraining different from test-time chain-of-thought?
- Can contrastive learning teach models to switch between logical and emotional reasoning?
- Do base models already contain latent behavioral principles waiting to be amplified?
- Why does adversarial training force deeper reasoning than surface imitation?
- What makes training-free approaches like Soft Thinking preferable to SoftCoT?
- Can auxiliary modules preserve reasoning without catastrophic forgetting?
- What does pass@k reveal about base model reasoning capacity?
- How much does pretraining contribute to ToM performance versus task-specific training?
- Can a single correct example seed exponential improvement in mathematical reasoning?
- What pretraining formats encode latent reasoning strategies that RLVR can surface?
- Can safety training and reasoning training be combined without losing calibration?
- Can approximate or noisy reference answers work for RL-based reasoning training?
- Why does critique training produce deeper understanding than imitation training?
- Can training-time debate work on tasks beyond mathematics and verifiable answers?
- Which domains need knowledge injection versus reasoning-focused training?
- Can testing prior knowledge and checking understanding improve explanation outcomes?
- Why does reasoning transfer across different numbers but factual recall does not?
- Why do recursive belief models require different training than logical derivation?
- How does factoring perception from reasoning improve sparse-label learning?
- Why does structured stochasticity help reasoning more than naive randomness?
- How does business logic specification replace annotated training datasets?
- What makes a model fail to activate relevant skills from its own harness?
- How much of Occamy's result comes from training versus the base model?
- How much reasoning catalyst data is actually needed for improvement?
- Can reasoning skills trained on law improve performance in STEM?