Line of inquiry
Inquiring lines›How should we train models for cap…›What systematic failures and vulne…›this line of inquiry
How do self-generated feedback mechanisms enable effective model learning?
A broader line of inquiry — a family of 42 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 42
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How do reward models and self-improvement mechanisms interact in training?
- How should training incorporate external critique versus encouraging self-correction?
- How much can externalized skills improve models before hitting diminishing returns?
- Does curriculum-based training keep small models perpetually at their learning edge?
- Does the generation-verification gap define where self-rewarding actually works?
- Does the productive difficulty band ever stabilize during training?
- At what capability level does the generation-verification gap make intrinsic rewards insufficient?
- Why does self-judgment of success or failure work without ground truth labels?
- How do self-play and human-anchored rewards separate competence from convention?
- What distinguishes surface mechanisms from the training regimes that produce them?
- How does adversarial collapse threaten unsupervised self-play skill construction?
- How do intrinsic motivation principles explain why generating novel challenges improves learning?
- How do different training objectives shift whether models over-predict or under-predict?
- Why does imitation learning alone plateau without outcome-based refinement?
- What training interventions could close the perception-action gap?
- Can a static evaluator become the performance ceiling for an improving actor?
- Can capability boundary collapse be reversed through external data?
- What makes a sub-goal verifiable enough to provide dense feedback signals?
- What training signals would models need to learn reciprocal common-ground construction?
- What other adaptive internal phenomena could signal system behavior improvements?
- How does temporal anchoring maintain the learning signal in self-rewarding loops?
- How does domain shift expose failures in fixed self-improvement mechanisms?
- Does self-conditioning improve belief-behavior alignment better than external priors?
- What emergent behaviors do models develop when trained on underspecified pedagogical tasks?
- Why do metric choices constrain which model capabilities get developed?
- How does a challenger's escalating difficulty function as curriculum?
- How do training objectives shape what a world model actually learns?
- Why does asymmetric self-play create naturally calibrated difficulty better than fixed curricula?
- Can models generate their own training curriculum during offline dreaming?
- How do developmental curriculums emerge from learning progress signals?
- What makes exploration a verifiable and measurable training objective?
- Why does optimizing only quality cause model collapse in self-improvement loops?
- What makes content informative and not-yet-mastered for reinforcement during pretraining?
- How does scaffolding unstable mechanics improve reinforcement learning for search?
- How does trajectory burstiness compare to other structural properties that shape emergent capabilities?
- What distinguishes intrinsic metacognition from extrinsic human-designed loops?
- What separates bootstrapping gains from sustained self-improvement gains?
- Does extended exoskeleton use eventually produce meaningful skill transfer?
- How does Goodhart's Law apply to proxy rewards in self-training systems?
- Can unified policies handle negative feedback and critique transformation simultaneously?
- What makes Effective Rank Acceleration a stable training signal for dual-channel incentives?
- What features does a sample reinforce when it moves bands?