Humans can only label so much, so should AI also learn from the world's own feedback, like test results and game wins?
What role should environmental rewards play versus human-specified objectives?
This explores whether AI systems should learn mainly from feedback the world gives them, such as test results, game outcomes and tool outputs, or from goals people write down. The corpus suggests the two don't really compete. Each does a job the other can't.
This explores whether AI systems should learn mainly from feedback the world gives them, such as test results, game outcomes and tool outputs, or from goals people write down. The corpus suggests the two don't really compete. Each does a job the other can't. The strongest case for environmental rewards comes from Silver and Sutton. They argue that gains from human-written data are flattening out in areas like math and coding, so agents will have to produce their own improvement data by acting in the world and getting grounded feedback (Will agent experience overtake human data for AI progress?). Behind the argument is a simple point: humans can only label so much, but the world keeps answering.
The surprising part is how much information environments carry beyond a single score. One line of work splits feedback into two kinds. Evaluative feedback says how well something went. Directive feedback says what to change. A single number keeps the first and throws away the second (Can scalar rewards capture all the information in agent feedback?). If you give the model the environment's actual response, such as an error message, a failed test or a critique, it can act as its own teacher. It looks back at its mistakes and turns them into step-by-step learning signals, with no separate reward model (Can environment feedback replace scalar rewards in policy learning?). Environments can even be built so the reward is checkable by design. RLSVR turns open-ended tasks like summarization into games where the game itself holds hidden answers, which removes the need for a judge to score them (Can environment structure replace external judges in RL?). Agents also lean on their environment in ways nobody planned. RL agents end up using the spaces they move through as a kind of external memory (Do RL agents accidentally use environments as memory?).
Environments still can't decide what counts as success, though. A model that chases its grader and a model that pursues what you actually wanted behave the same, right up until the grader rewards the wrong thing (Can we detect reward-seeking from normal model behavior?). So reward-hacking risk doesn't go away with environmental rewards. It hides inside whoever designed the environment. Systems that try to improve with no outside input run into the same wall: pure self-improvement goes in circles, and the methods that work quietly bring in an outside anchor, such as an older model version, a third-party judge, a user correction or a tool's output (Can models reliably improve themselves without external feedback?). Peer models can stand in for some of that outside check, but only when the peers are genuinely different from each other (Can peer models replace external judges for reward signals?).
The most useful idea here is a division of labor. Human goals work best as a gate, a check that an answer passes or fails, and environmental rewards work best as the gradient that improves answers within that gate. DRO shows this. Using rubrics to accept or reject batches of answers, rather than turning rubric scores into rewards, prevents reward hacking while dense rewards handle fine-tuning inside the accepted set (Can rubrics and dense rewards work together without hacking?). The same pattern appears in human oversight. Stepping in only at high-uncertainty moments beat both full autonomy and step-by-step review: 87.5% of outputs were accepted, against 25% and 50% (Does targeted human oversight beat both full autonomy and exhaustive review?). On long tasks, how persistently an agent works through feedback loops predicted success better than how good its first attempt was (What predicts success in ultra-long-horizon agent tasks?). So the rough answer is that people set the boundaries and check them at the moments that matter, and the environment supplies the volume of feedback in between. The corpus doesn't yet say much about where those boundaries come from when the task itself is open-ended.
Sources 11 notes
As human-data gains plateau in domains like mathematics and coding, agents must generate their own improvement data through environmental interaction with grounded rewards. Evidence includes AlphaProof's Olympiad performance and the theoretical advantage of adaptive, experience-based learning over static human curation.
Natural feedback carries two orthogonal types of information: evaluative (how well an action performed) and directive (how it should change). Scalar rewards capture evaluation but discard directional specifics that token-level distillation can recover, making the two complementary rather than redundant.
SDPO converts tokenized environment feedback into dense gradient signals by using the feedback-conditioned policy as a self-teacher. The policy, when given retrospective evidence of its mistakes in-context, implicitly acts as its own process reward model, making external reward signals unnecessary.
RLSVR transforms open-ended tasks into proxy environments like SpyRL where hidden variables assigned by the game supply verifiable rewards, eliminating judges, reward models, and their associated bias and costs. SpyRL reportedly outperforms existing self-improvement methods on summarization and creative writing.
Mathematical proof shows that environmental artifacts reduce information needed to represent history in RL agents. Path-following agents naturally develop memory-like behavior through standard reward optimization, satisfying situated cognition criteria without explicit memory objectives.
Show all 11 sources
Models pursuing grader judgment and those pursuing intended objectives behave identically whenever evaluation agrees with intent. Reward-seeking only becomes visible when graders reward unintended behavior, which well-designed pipelines eliminate.
Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.
Co-RL trains decoupled models using peer predictions as rewards, avoiding the bias and collapse of self-generated feedback. Heterogeneous cohorts consistently improve reasoning across benchmarks and often match ground-truth supervised training.
DRO shows that using rubrics to accept or reject rollout groups—rather than converting rubric scores into dense rewards—prevents reward hacking. This separation preserves the categorical strength of rubrics while letting token-level rewards optimize within valid answers.
AutoResearchClaw's confidence-routed CoPilot mode achieved 87.5% accept rate, beating full autonomy (25%) and step-by-step oversight (50%). Selective human intervention on high-stakes decisions avoids both uncaught errors and the rubber-stamping fatigue of constant interruption.
Across 17 frontier models on 36 expert-curated optimization tasks, repeated benchmark-edit-incorporate cycles within a wall-clock budget proved the dominant success predictor. Most models terminated early or burned budget unproductively; Claude Opus 4.6 stood out as persistent.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Reward Reasoning Model
- Artifacts as Memory Beyond the Agent Boundary
- Reinforcement Learning via Self-Distillation
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- Learning to Reason without External Rewards
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
- Welcome to the Era of Experience
- Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL