Line of inquiry
Inquiring lines›How do training signals reliably a…›What training signals and data cur…›this line of inquiry
Does reinforcement learning create new reasoning capabilities or optimize existing ones?
A broader line of inquiry — a family of 104 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 104
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can reinforcement learning add new capabilities or only remove inaccurate knowledge?
- Can RL create new reasoning primitives that pretraining never established?
- Does RL amplify existing reasoning or create genuinely new computational strategies?
- How does RL refine reasoning paths without simply adding model capability?
- Does RL primarily teach when to use reasoning or how to reason?
- When does RL discover genuinely novel reasoning strategies versus timing optimization?
- Does RL refine existing knowledge or discover entirely new capabilities?
- Does RL teach models new reasoning or just better timing?
- When does reinforcement learning actually produce true reasoning gains in models?
- Does RL teach models when to use reasoning or how to reason?
- Does reinforcement learning preserve reasoning quality better than supervised fine-tuning?
- Can extended RL training unlock genuinely new reasoning strategies models cannot discover otherwise?
- Why does prolonged RL discover strategies absent from any base model sample?
- Does reinforcement learning learn optimal per-turn reasoning discipline?
- How does RL compress reasoning path diversity during training?
- Can reinforcement learning fix the reasoning gaps that supervised fine-tuning misses?
- Can base models spontaneously produce reasoning traces without any RL training?
- Why do reasoning gains from RL require models trained with headroom and edge-of-competence data?
- Can RLVR expand a model's reasoning capabilities beyond its training ceiling?
- Can RL teach when to use reasoning versus when to respond directly?
- Why do current RLVR methods fail to expand reasoning capability beyond base model boundaries?
- Why does RL improve sampling efficiency but not expand capability boundaries?
- Can single-problem fine-tuning match full RL pipeline reasoning gains?
- Can reinforcement learning add missing domain knowledge to fine-tuned reasoning models?
- Why does RL behavior differ between standard reasoning tasks and complex planning domains?
- Can reinforcement learning teach AI when to ask clarifying questions?
- Can one training example activate mathematical reasoning in RL-trained models?
- Can reinforcement learning improve how accurately models explain themselves?
- Can reinforcement learning close the gap between LLM reasoning and action?
- Does format-based pretraining determine how models respond to reinforcement learning?
- Why do overtrained domains show different RL training outcomes than novel tasks?
- Are RLVR models worse than non-reasoning models for subjective annotation?
- How do sparse parameter updates enable when-not-how training to work?
- Why does outcome-based RL specifically lose diversity during training?
- Can out-of-distribution tests expose memorization in reinforcement learning fine-tuned models?
- Does RLVR teach new reasoning or activate existing pretraining capabilities?
- Can verifier-free RL work without manual preference labels or task-specific training?
- Does reinforcement learning teach models how to reason or when to reason?
- What limits RL's ability to scale for reasoning at training time?
- Does RLVR expand model capability or reorganize existing capability?
- How does reinforcement learning on outcomes reinforce template-matching rather than computation?
- Why does standard RL cause traces to collapse into redundant reasoning paths?
- What role does reinforcement learning play in optimizing inference compute?
- What limits RLVR effectiveness beyond mathematical and coding domains?
- Can smaller models achieve domain expertise through focused RL training?
- Why does supervised fine-tuning on diverse demonstrations expand exploration diversity compared to RL?
- How do extrapolative and contextual generalization measure RL reasoning gains?
- How does imitation pretraining followed by RL exploration compare to either method alone?
- How does baseline capability level affect RL improvement ceiling?
- Can in-context learning replicate the timing effects that RL teaches models?
- What capacity threshold determines whether RL teaches activation versus shortcut learning?
- What role does natural language play in breaking reinforcement learning performance plateaus?
- What makes supervised fine-tuning worsen RL exploration later?
- How does pretraining determine what RL can later teach a model?
- Can combining SRL with RLVR outperform either method used alone?
- Can in-context reinforcement learning match human sample efficiency on real problems?
- Why does the pretrained prior determine the exploration ceiling?
- What breaks when you apply reinforcement learning after supervised fine-tuning?
- How does reinforcement learning compare to differentiable joint training for RAG?
- Why does extended reasoning training improve exploration without adding new capabilities?
- How does post-training shift models from passive prediction to on-policy action?
- Why does exploration diversity behave differently under reinforcement learning versus supervised fine-tuning?
- What does RL post-training actually teach reasoning systems?
- Does the pretrained model prior limit RL search capability more than the optimization algorithm itself?
- Why does multi-turn RL generate orders of magnitude more tokens than single-turn?
- What scaling properties emerge from RL training dynamics beyond verification?
- Can RL directly optimize attention distributions instead of text generation?
- Why do models follow a two-phase pattern of procedural then strategic learning?
- What training duration is actually needed for RL to expand capabilities?
- Why does reinforcement learning training degrade model calibration?
- Can LLM-synthesized behavioral heuristics compete with learned policy improvements?
- How should humans specify deterministic abstractions of RL problems?
- What distinguishes high-signal prompts from low-signal ones in RL training?
- How does LLM simulation of APIs avoid instability without sacrificing training signal?
- Why does medium difficulty outperform both easy and hard RLVR training samples?
- How does reinforcement learning differ from chain-of-thought distillation?
- What distinguishes RL that creates new capabilities from RL that merely teaches timing?
- How does Supervised RL bridge the gap between SFT and RLVR?
- Does RL training redirect self-doubt into productive gap analysis?
- How can verifier-free reinforcement learning handle reasoning without task-specific checks?
- Can models learn both what and how to study through reinforcement learning?
- Does sparsity in RL arise from training on policy-distribution data?
- How do verifier-free and adversarial approaches compare in extending reasoning RL?
- Why does combining reasoning distillation with RLVR outperform either training stage alone?
- Does task ordering affect multi-task reinforcement learning outcomes?
- How does non-reasoning SFT prevent overfitting before RL training begins?
- How does prolonged RL training differ from standard RLVR approaches?
- How do thought actions represent policy improvement steps in practice?
- Can one training example activate mathematical reasoning without reinforcement learning?
- How do verifier-free RL patterns differ from traditional RLHF approaches?
- How does trajectory filtering handle noise when language models use code execution tools?
- Why does early experience provide better warm-starts for downstream reinforcement learning?
- How does the pretrained prior constrain the ceiling for empathy RL improvements?
- How do residual connections and layer norm stabilize training in deep RL?
- Can the exploration ceiling be raised beyond what pretraining established?
- How do RL training and base models differ in creating MI peaks?
- Which recipe choices determine the asymptotic ceiling in RL training?
- Can approximate or noisy reference answers work for RL-based reasoning training?
- What makes software engineering environments better suited for RL than other interactive domains?
- Can meta-reinforcement learning explain why this bias pattern emerges rationally?
- Why does online RL succeed where supervised training fails for self-correction?
- What makes some tasks bounded enough for reliable RL?
- How does active selection of training content differ from random reinforcement sampling?
- How does behavior cloning reduce complexity before RL training in rerankers?