Line of inquiry
Inquiring lines›How do training and design choices…›What determines whether training i…›this line of inquiry
Can self-generated feedback reliably guide model training without ground truth?
A broader line of inquiry — a family of 63 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 63
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why does self-correction during generation produce reliable labels without exemplars?
- Why does self-consistency fail as a proxy reward for correctness?
- How does error avalanching compound failures in self-training iterations?
- Can self-consistency checks fully prevent error avalanching in self-training loops?
- What role does the self-consistency threshold play in preventing error reinforcement?
- Can applicability conditions and veto rules make self-training stable across substrates?
- How should training incorporate external critique versus encouraging self-correction?
- Why does self-judgment of success or failure work without ground truth labels?
- Does the generation-verification gap define where self-rewarding actually works?
- How do prior errors in context history amplify future failures over time?
- What makes self-consistency a sufficient training target for the judge role?
- How does self-consistency as a proxy reward incentivize confident-but-wrong answers?
- How does self-improvement capability vary across memory, retrieval, and update tasks?
- How does adversarial collapse threaten unsupervised self-play skill construction?
- Can self-training drift be prevented by applying student compatibility filtering?
- Can AI-generated explanations of errors teach as effectively as self-resolution?
- Can population diversity in self-improvement prevent error avalanching failures?
- How does self-consistency compare to confidence as a proxy reward signal?
- What external anchors prevent self-editing from collapsing into circularity?
- Can a model evaluate its own improvements without degrading over iterations?
- Why do error avalanches accelerate in self-training loops without verification?
- Why does every reliable LLM self-improvement require external intervention or verification?
- What failure modes emerge when model-generated content trains on itself iteratively?
- Can multiple verification approaches together overcome the self-improvement ceiling?
- How does an external evaluation anchor prevent self-improvement from becoming circular?
- Why does self-generated training data outperform externally curated domain examples?
- How do prior errors in context history amplify future mistakes in long tasks?
- Can synthetic self-play data teach models when to disagree?
- What external signals make self-improvement loops bounded rather than circular?
- Why does self-generated training data outperform externally sourced data?
- Does self-conditioning improve belief-behavior alignment better than external priors?
- Why does systematic overconfidence on self-generated outputs compound autoregressive errors?
- What collapse dynamics constrain recursive self-improvement in current evidence?
- Why does uncontrolled self-revision drift toward instance-specific overfitting?
- Do self-generated explanations outperform passive instruction for oversight?
- Why do models trained on critique fail at self-critique despite strong other-model evaluation?
- How does distribution mismatch between training and deployment break self-correction?
- How does domain shift expose failures in fixed self-improvement mechanisms?
- Why do current metacognitive training loops fail when agents encounter new domains?
- Why do method-level improvements avoid the generation-verification gap that parameter-level improvements face?
- How do instruction backtranslation and MAGPIE demonstrate self-generation principles?
- Why does research-direction judgment validation limit fully closed self-improvement?
- Does the pretrained prior actually constrain what internalized search can discover?
- How does temporal anchoring maintain the learning signal in self-rewarding loops?
- Can held-out validation gates prevent optimizer hallucinations in skill proposals?
- Why do self-consistency methods fail where pretraining bias is strongest?
- How does adversarial self-play during training differ from single-model self-revision?
- Can capability boundary collapse be reversed through external data?
- How do test harnesses guide reflection better than transcripts alone?
- Why does optimizing only quality cause model collapse in self-improvement loops?
- Why does filtering for correct examples prevent error compounding in self-training?
- Can trustworthy scoring prevent persistent iteration from compounding errors?
- Why does online RL succeed where supervised training fails for self-correction?
- Why do unchecked self-edits accumulate drift toward overfitting or incoherence?
- How does self-observation enable experts to verify their own judgment?
- What makes deliberate practice on your own errors more effective than copying others?
- Why does asymmetric self-play create naturally calibrated difficulty better than fixed curricula?
- What distinguishes intrinsic metacognition from extrinsic human-designed loops?
- When does provable stability in latent dynamics fail to preserve fidelity?
- How does Goodhart's Law apply to proxy rewards in self-training systems?
- What separates bootstrapping gains from sustained self-improvement gains?
- What makes recursive self-improvement circular or well-founded?
- Can predictive self-supervision work on unlabeled sequential visual data?