Can language models design better reward functions than humans?
Can GPT-4 write and automatically refine reward code to outperform manual reward engineering? This explores whether LLMs can solve a core bottleneck in reinforcement learning—specifying what an agent should optimize for.
EUREKA, introduced in this paper, has a coding LLM (GPT-4) write executable reward-function code directly from an environment's source, then evolves that code through iterative search. Tested across "29 open-source RL environments that include 10 distinct robot morphologies," EUREKA "outperforms human experts on 83% of the tasks" and achieves "an average normalized improvement of 52%" — without "any task-specific prompting or pre-defined reward templates." The headline capability result is dexterity: combined with curriculum learning, EUREKA-designed rewards let a simulated Shadow Hand perform pen-spinning maneuvers, a skill the paper says manual reward engineering had not achieved.
The paper names three components. "Environment as context" feeds the raw environment source code (the state and action variables, not simulator internals) to the LLM so it can zero-shot a plausible reward function without domain-specific prompting. "Evolutionary search" then samples K=16 reward candidates per iteration, keeps the best, and mutates it for the next round — five iterations, five independent runs per environment — exploiting the fact that "the probability that all reward functions from an iteration are buggy exponentially decreases as the number of samples increases." The improvement signal across iterations is "reward reflection": a textual summary of policy-training statistics — the numeric trajectory of each named reward component plus the task fitness score — fed back to the LLM so it can diagnose which sub-term of its own code to edit. The paper is explicit that the scalar fitness score alone "lacks in credit assignment," which is why reflection decomposes feedback down to the reward's own named components.
This is a different point in the automated-reward-design space than Can LLMs design reward functions for reinforcement learning?: MEDIC has the LLM solve a simplified planning problem and converts the resulting policy into a shaping term, whereas EUREKA has the LLM write the reward function's code directly and iterates on that code using execution and training telemetry as the editing signal — no abstraction-solving step, no guide policy. Reward reflection is also a cheaper route to the dense, step-level feedback that Why do outcome-based reward models fail at intermediate step evaluation? says normally requires expensive annotation: EUREKA gets per-component granularity for free because the reward code already exposes its own named sub-terms, so the training run self-reports instead of a human labeling steps — a fourth variant alongside the three Can trajectory structure replace hand-annotated process rewards? describes. And where Why is objective design the real bottleneck in AI discovery? argues the real bottleneck in open-ended discovery is the creativity of specifying what to optimize, EUREKA shows that for a bounded class of RL environments an LLM can take over the downstream half of that problem — turning an existing task metric into working reward code — while still depending on a human-specified fitness function to search against.
The excerpt measures performance only on existing simulated RL benchmarks with an available task fitness function; it does not show EUREKA generating the fitness function itself, nor does it report results outside simulation or on tasks lacking a numeric ground-truth metric to search against. The RLHF extension is validated with a 20-person preference study on a single locomotion task, not a general claim about aligning reward code to open-ended human intent. The paper's closing claim — that combining LLMs with evolutionary algorithms is "a general and scalable approach" to open-ended search — is stated as a conjecture about applicability beyond reward design, not a result this paper itself demonstrates.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do reward signal properties affect model reasoning and safety? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can LLMs design reward functions for reinforcement learning?
Can language models help automate the notoriously difficult task of designing reward shaping functions for sparse-reward RL, and if so, how might we structure that collaboration to work around LLMs' weaknesses in stochastic control?
different mechanism: EUREKA writes reward code directly and edits it, no abstraction-solving step or guide policy
-
Why do outcome-based reward models fail at intermediate step evaluation?
Outcome-based reward models (ORMs) evaluate only final results, creating a mismatch with the need to assess reasoning quality at intermediate steps. Understanding this failure mode matters for building better AI reasoning systems.
reward reflection gets dense per-component feedback for free since the reward code exposes its own named sub-terms
-
Can trajectory structure replace hand-annotated process rewards?
Recent methods extract step-level supervision directly from how agent trajectories are structured—trees, expert alignments, tool calls—rather than training separate reward models. Can this structural approach consistently avoid annotation costs?
reward reflection is a fourth instance: dense feedback derived from the reward code's own structure rather than annotation
-
Why is objective design the real bottleneck in AI discovery?
If AI agents can search hypothesis spaces efficiently, what makes defining the right objective function harder than finding solutions? This explores whether creativity in science lies more in problem formulation than problem-solving.
EUREKA automates the downstream half of reward design but still needs a human-specified fitness function to search against
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Eureka: Human-Level Reward Design via Coding Large Language Models
- Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
- Efficient Reinforcement Learning via Large Language Model-based Search
- Reward Reasoning Model
- Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
- Reward-Robust RLHF in LLMs
- Learning to Reason without External Rewards
- Self-Rewarding Language Models
Original note title
LLM-generated reward code, refined through evolutionary search and reward reflection, outperforms human experts on 83% of RL tasks