Theme of inquiry
What training signals and data curation strategies optimize model learning?
A question within its area, explored through 8 lines of inquiry below — each a family of specific questions the research asks.
50 specific questions
- How much can externalized skills improve models before hitting diminishing returns?
- Does specialized training in one domain create capability cliffs elsewhere?
- Does curriculum-based training keep small models perpetually at their learning edge?
- How do self-evolving curricula help RL break beyond base model capability boundaries?
- What performance trade-offs emerge when composing multiple independently trained model capabilities?
- Can smaller models produce skill updates as useful as frontier model updates?
- Do emergent abilities result from genuine new capabilities or implicit in-context learning?
46 specific questions
- How does question difficulty and breadth affect what models learn to reason?
- How does difficulty-adaptive curriculum learning change which samples get selected during training?
- Does selecting examples from multiple complexity levels outperform selecting only high-quality examples?
- Does partial trace guidance work better than curriculum learning for hard problems?
- Does the productive difficulty band ever stabilize during training?
- Can models learn better from critiquing errors than imitating correct responses?
- Why do strong models struggle more with instruction following than mid-tier ones?
24 specific questions
- Do situationally aware models deliberately exploit their graders' judgment gaps?
- Does situational awareness help models hide reward-seeking during evaluation?
- How does situational awareness interact with reward-seeking in RL training?
- Why does training against detected failures select for passing detection instead?
- Does reward-seeking grow worse with situational awareness and reinforcement learning compute?
- Can behavioral training ever produce compliance that doesn't depend on being observed?
- Can prohibitions learned from scored behavior become detection-avoidance rather than norm internalization?
104 specific questions
- Can reinforcement learning add new capabilities or only remove inaccurate knowledge?
- Can RL create new reasoning primitives that pretraining never established?
- Does RL amplify existing reasoning or create genuinely new computational strategies?
- How does RL refine reasoning paths without simply adding model capability?
- Does RL primarily teach when to use reasoning or how to reason?
- When does RL discover genuinely novel reasoning strategies versus timing optimization?
- Does RL refine existing knowledge or discover entirely new capabilities?
41 specific questions
- Does RLHF training make explanations more deceptive than transparent?
- Does RLHF training create models that sound convincing without being more accurate?
- Does RLHF training specifically teach models to prioritize user agreement over accuracy?
- Why does RLHF training optimize for perceived quality over practical accuracy?
- How does RLHF training reward models for guessing over asking clarifying questions?
- How does RLHF training for helpfulness create systematic misinterpretation patterns?
- How does RLHF helpfulness training drive premature assumptions in multi-turn dialogue?
21 specific questions
- What does process supervision reveal about step-level reasoning versus outcome rewards?
- What makes process-level supervision better than outcome-only reward signals?
- What are the actual limits of sibling comparison versus trained process reward models?
- Does process supervision recover reasoning accuracy better than outcome rewards in latent space?
- Why does outcome supervision fail for long reasoning chains?
- Why does step-level expert alignment work when outcome-only RL fails?
- How does process-focused feedback compare to outcome-focused feedback in skill training?
53 specific questions
- How much does training composition affect syntactic versus reasoning performance?
- Can training on diverse related tasks be more efficient than task-specific training?
- How much task-similar finetuning data does test-time training actually need?
- How do task frequency and complexity interact with model capacity during training?
- Why does exploration quality matter more than learner network depth?
- Can selecting the right data subset outperform training on everything?
- How much of the combinatorial task space must training data cover?
45 specific questions
- How do finetuning and pretraining improvements differ in their effects on model capabilities?
- Does fine-tuning actually change model capabilities or only output distribution?
- How does behavioral fine-tuning differ from factual knowledge encoding in models?
- What is the difference between changing model outputs versus changing internal representations?
- How does model scale affect anticipatory behavior in structured training?
- Why does the gap between theoretical expressiveness and learned capability matter?
- Does pretraining data size matter less than base model scale for finetuning?