Can an AI write its own grading rubric for a robot's training, well enough to beat a human-written one?
What class of RL problems can LLMs reliably turn into working reward code?
This explores which kinds of reinforcement learning tasks an LLM can reliably write the reward function for (the code that scores how well an agent is doing), and where that ability runs out.
This explores which kinds of RL problems an LLM can reliably turn into working reward code, and where that ability stops. The short answer from the corpus: LLMs do well when the task can be simulated, when success can be measured and fed back to them, and when the problem can be reduced to a simpler version they can reason about. The collection doesn't offer a clean map of the boundary. It shows two working recipes and several warning signs.
The headline result is EUREKA Can language models design better reward functions than humans?. GPT-4 writes reward code for simulated robotics and control tasks, and its rewards beat human-written ones on 83% of 29 benchmarks, including dexterous hand-manipulation skills that humans had trouble specifying. The important detail is that GPT-4 doesn't get it right in one shot. It writes candidates, the system trains agents on them, and statistics from those training runs come back as feedback for the next round of edits. So the class of problems is really "tasks you can cheaply run in a simulator many times." The reliability comes from the loop, not from the first draft. MEDIC Can LLMs design reward functions for reinforcement learning? reaches a similar result by a different route. The LLM solves a simplified, deterministic version of a messy, randomized task, turns that plan into intermediate rewards for the real problem, and a separate model-based checker vets the output before it's used. Both recipes share the same lesson: an LLM is a good reward author when something outside the LLM checks its work.
The edges show up in nearby notes. When a task is mostly about satisfying many interacting constraints, LLMs stall at about 55–60% constraint satisfaction no matter how large the model is Do larger language models solve constrained optimization better?. That's a reason to doubt rewards that must encode many hard requirements at once. Even a simple, correct-looking reward can teach the wrong thing. Plain right/wrong rewards push models toward confident guessing, and adding a scoring term that penalizes confident wrong answers fixes this Does binary reward training hurt model calibration?. Reward code that compiles and runs is not the same as reward code that shapes the behavior you wanted.
Other notes take a different approach to the same goal. Instead of writing reward code, the LLM produces the reward signal directly. Tree search can rank solution paths and replace human labels Can tree search replace human feedback in LLM training?. LLM judges anchored to reference answers can supervise training about as well as a trained reward model Can reference examples make LLM judges reliable enough for self-improvement?. LLMs can even stand in for the environment, simulating a search engine Can LLMs replace search engines during agent training? or a user. But simulated users drift from their own goals and corrupt the training signal unless their goals are tracked explicitly Why do LLM user simulators fail to track their own goals?.
The takeaway: the corpus's reliable zone is simulated physical and control tasks where rewards can be tested by running them. For fuzzier domains like dialogue, open-ended reasoning, and heavily constrained planning, the field has mostly moved away from LLM-written reward code toward LLM-as-judge and LLM-as-environment, each with its own failure modes. The corpus doesn't directly test LLM-authored reward code on those fuzzier domains, so that boundary is an inference, not a measured result.
Sources 8 notes
EUREKA uses evolutionary search and training-statistic feedback to evolve reward functions written by GPT-4, outperforming human-designed rewards on 83% of 29 RL benchmarks with 52% average improvement, including novel dexterous manipulation skills.
MEDIC shows that LLMs can generate effective reward shaping functions by first solving a deterministic, simplified version of the RL problem, then converting the resulting plan into shaping rewards for the original stochastic task. A model-based critic validates LLM outputs before deployment.
Across constrained-optimization tasks, LLMs converge to ~55–60% constraint satisfaction independent of architecture, parameter count, or training regime. Reasoning models do not systematically outperform standard models, suggesting a fundamental ceiling rather than a scaling gap.
Binary correctness rewards incentivize high-confidence guessing because they don't penalize confident wrong answers. Adding the Brier score as a second reward term mathematically guarantees joint optimization of accuracy and calibration without trade-off.
AlphaLLM uses tree search outcomes and three critic models to derive dense reward signals equivalent to human-labeled feedback. Tree structure naturally ranks solution paths by success, replacing the annotation oracle that standard RLHF requires.
Show all 8 sources
Anchoring LLM-judges to reference answers with explicit usage instructions improved judge accuracy by 6.8% and enabled self-improvement training via DPO to match finetuned reward model performance on AlpacaEval and Arena-Hard benchmarks.
ZeroSearch and SSRL demonstrate that LLMs can generate relevant documents and search results from internal knowledge, with 14B simulators matching or exceeding real search engines. Curriculum degradation and test-time scaling optimize this approach for training without API costs.
The UGST framework breaks user goals into profile, policy, task, requirements, and preferences—each with explicit status tracking. A three-stage method (steering, SFT, GRPO) progressively internalizes goal alignment, reducing the misalignment that corrupts RL training signals.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Eureka: Human-Level Reward Design via Coding Large Language Models
- References Improve LLM Alignment in Non-Verifiable Domains
- Efficient Reinforcement Learning via Large Language Model-based Search
- Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
- Self-Improving Model Steering
- Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling
- Reward Reasoning Model
- Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge