Line of inquiry
Inquiring lines›How can we ensure training objecti…›How do reward models and preferenc…›this line of inquiry
Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones?
A broader line of inquiry — a family of 100 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 100
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does RL amplify existing reasoning or create genuinely new computational strategies?
- Can reinforcement learning add new capabilities or only remove inaccurate knowledge?
- How does RL refine reasoning paths without simply adding model capability?
- Does RL primarily teach when to use reasoning or how to reason?
- Can RL create new reasoning primitives that pretraining never established?
- When does RL discover genuinely novel reasoning strategies versus timing optimization?
- When does reinforcement learning actually produce true reasoning gains in models?
- Does RL teach models when to use reasoning or how to reason?
- Does reinforcement learning preserve reasoning quality better than supervised fine-tuning?
- Does RL teach models new reasoning or just better timing?
- Does RL refine existing knowledge or discover entirely new capabilities?
- Can extended RL training unlock genuinely new reasoning strategies models cannot discover otherwise?
- Does reinforcement learning learn optimal per-turn reasoning discipline?
- Why does prolonged RL discover strategies absent from any base model sample?
- Can reinforcement learning fix the reasoning gaps that supervised fine-tuning misses?
- How does RL compress reasoning path diversity during training?
- Can base models spontaneously produce reasoning traces without any RL training?
- Can RLVR expand a model's reasoning capabilities beyond its training ceiling?
- Why do reasoning gains from RL require models trained with headroom and edge-of-competence data?
- Can single-problem fine-tuning match full RL pipeline reasoning gains?
- Can RL teach when to use reasoning versus when to respond directly?
- Why do current RLVR methods fail to expand reasoning capability beyond base model boundaries?
- Why does RL behavior differ between standard reasoning tasks and complex planning domains?
- Can reinforcement learning add missing domain knowledge to fine-tuned reasoning models?
- Why does RL improve sampling efficiency but not expand capability boundaries?
- Can reinforcement learning teach AI when to ask clarifying questions?
- Can reinforcement learning improve how accurately models explain themselves?
- Can one training example activate mathematical reasoning in RL-trained models?
- Can reinforcement learning close the gap between LLM reasoning and action?
- Are RLVR models worse than non-reasoning models for subjective annotation?
- Does format-based pretraining determine how models respond to reinforcement learning?
- How does pretraining quality versus quantity affect downstream RL gains?
- Why do overtrained domains show different RL training outcomes than novel tasks?
- Can verifier-free RL work without manual preference labels or task-specific training?
- How do sparse parameter updates enable when-not-how training to work?
- Can out-of-distribution tests expose memorization in reinforcement learning fine-tuned models?
- Does reinforcement learning require sufficient pretraining to be effective?
- How does reinforcement learning on outcomes reinforce template-matching rather than computation?
- Why does outcome-based RL specifically lose diversity during training?
- Does reinforcement learning teach models how to reason or when to reason?
- What role does reinforcement learning play in optimizing inference compute?
- What limits RLVR effectiveness beyond mathematical and coding domains?
- Does RLVR expand model capability or reorganize existing capability?
- Does RLVR teach new reasoning or activate existing pretraining capabilities?
- Why does standard RL cause traces to collapse into redundant reasoning paths?
- Can explicitly optimizing for semantic diversity during RL training improve both quality and variation?
- Can RL training on small verifiable tasks transfer to real-world AI research?
- Can smaller models achieve domain expertise through focused RL training?
- Does the reinforcement learning improvement rate depend on model initialization?
- Should larger compute budgets allocate more resources to reinforcement learning?
- Can in-context reinforcement learning match human sample efficiency on real problems?
- Why does supervised fine-tuning on diverse demonstrations expand exploration diversity compared to RL?
- Can in-context learning replicate the timing effects that RL teaches models?
- What role does natural language play in breaking reinforcement learning performance plateaus?
- How do extrapolative and contextual generalization measure RL reasoning gains?
- How do Q-value models improve action selection compared to value models?
- How does imitation pretraining followed by RL exploration compare to either method alone?
- How does reinforcement learning compare to differentiable joint training for RAG?
- What hard-to-verify tasks will remain resistant to reinforcement learning?
- Can safety training reduce or eliminate metagaming after capabilities training?
- Can combining SRL with RLVR outperform either method used alone?
- Can negative reinforcement alone match full RL performance on domain tasks?
- Can RL directly optimize attention distributions instead of text generation?
- What breaks when you apply reinforcement learning after supervised fine-tuning?
- Why does exploration diversity behave differently under reinforcement learning versus supervised fine-tuning?
- What makes supervised fine-tuning worsen RL exploration later?
- Can pretraining alone achieve chess performance without reinforcement learning?
- Can LLM-synthesized behavioral heuristics compete with learned policy improvements?
- Why does multi-turn RL generate orders of magnitude more tokens than single-turn?
- Why do models follow a two-phase pattern of procedural then strategic learning?
- Why does reinforcement learning training degrade model calibration?
- How can verifier-free reinforcement learning handle reasoning without task-specific checks?
- What does RL post-training actually teach reasoning systems?
- How does LLM simulation of APIs avoid instability without sacrificing training signal?
- Why does medium difficulty outperform both easy and hard RLVR training samples?
- How should humans specify deterministic abstractions of RL problems?
- What distinguishes high-signal prompts from low-signal ones in RL training?
- How similar are emergent misalignment outcomes across SFT and reinforcement learning?
- How does reinforcement learning differ from chain-of-thought distillation?
- Does RL training redirect self-doubt into productive gap analysis?
- Can models learn both what and how to study through reinforcement learning?
- Does task ordering affect multi-task reinforcement learning outcomes?
- Does sparsity in RL arise from training on policy-distribution data?
- How do verifier-free and adversarial approaches compare in extending reasoning RL?
- How does Supervised RL bridge the gap between SFT and RLVR?
- What behavioral changes occur during reward learning training?
- How do thought actions represent policy improvement steps in practice?
- How do verifier-free RL patterns differ from traditional RLHF approaches?
- What role does online RL play in scaling GUI agents?
- How does prolonged RL training differ from standard RLVR approaches?
- How does non-reasoning SFT prevent overfitting before RL training begins?
- Can one training example activate mathematical reasoning without reinforcement learning?
- What class of RL problems can LLMs reliably turn into working reward code?
- How do residual connections and layer norm stabilize training in deep RL?
- What makes software engineering environments better suited for RL than other interactive domains?
- Should user simulators be trained via RL like agents or decomposed into trackable state components?
- How does RLSVR differ from using model probability or self-judgment?
- What makes a task at the edge of competence optimal for RL?
- What makes some tasks bounded enough for reliable RL?
- Can meta-reinforcement learning explain why this bias pattern emerges rationally?