SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

Can language models design better reward functions than humans?

Can GPT-4 write and automatically refine reward code to outperform manual reward engineering? This explores whether LLMs can solve a core bottleneck in reinforcement learning—specifying what an agent should optimize for.

Synthesis note · 2026-10-08 · sourced from Frontier AI Risk & RSI

EUREKA, introduced in this paper, has a coding LLM (GPT-4) write executable reward-function code directly from an environment's source, then evolves that code through iterative search. Tested across "29 open-source RL environments that include 10 distinct robot morphologies," EUREKA "outperforms human experts on 83% of the tasks" and achieves "an average normalized improvement of 52%" — without "any task-specific prompting or pre-defined reward templates." The headline capability result is dexterity: combined with curriculum learning, EUREKA-designed rewards let a simulated Shadow Hand perform pen-spinning maneuvers, a skill the paper says manual reward engineering had not achieved.

The paper names three components. "Environment as context" feeds the raw environment source code (the state and action variables, not simulator internals) to the LLM so it can zero-shot a plausible reward function without domain-specific prompting. "Evolutionary search" then samples K=16 reward candidates per iteration, keeps the best, and mutates it for the next round — five iterations, five independent runs per environment — exploiting the fact that "the probability that all reward functions from an iteration are buggy exponentially decreases as the number of samples increases." The improvement signal across iterations is "reward reflection": a textual summary of policy-training statistics — the numeric trajectory of each named reward component plus the task fitness score — fed back to the LLM so it can diagnose which sub-term of its own code to edit. The paper is explicit that the scalar fitness score alone "lacks in credit assignment," which is why reflection decomposes feedback down to the reward's own named components.

This is a different point in the automated-reward-design space than Can LLMs design reward functions for reinforcement learning?: MEDIC has the LLM solve a simplified planning problem and converts the resulting policy into a shaping term, whereas EUREKA has the LLM write the reward function's code directly and iterates on that code using execution and training telemetry as the editing signal — no abstraction-solving step, no guide policy. Reward reflection is also a cheaper route to the dense, step-level feedback that Why do outcome-based reward models fail at intermediate step evaluation? says normally requires expensive annotation: EUREKA gets per-component granularity for free because the reward code already exposes its own named sub-terms, so the training run self-reports instead of a human labeling steps — a fourth variant alongside the three Can trajectory structure replace hand-annotated process rewards? describes. And where Why is objective design the real bottleneck in AI discovery? argues the real bottleneck in open-ended discovery is the creativity of specifying what to optimize, EUREKA shows that for a bounded class of RL environments an LLM can take over the downstream half of that problem — turning an existing task metric into working reward code — while still depending on a human-specified fitness function to search against.

The excerpt measures performance only on existing simulated RL benchmarks with an available task fitness function; it does not show EUREKA generating the fitness function itself, nor does it report results outside simulation or on tasks lacking a numeric ground-truth metric to search against. The RLHF extension is validated with a 20-person preference study on a single locomotion task, not a general claim about aligning reward code to open-ended human intent. The paper's closing claim — that combining LLMs with evolutionary algorithms is "a general and scalable approach" to open-ended search — is stated as a conjecture about applicability beyond reward design, not a result this paper itself demonstrates.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do reward signal properties affect model reasoning and safety? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 136 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLM-generated reward code, refined through evolutionary search and reward reflection, outperforms human experts on 83% of RL tasks