RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Paper · arXiv 2411.15114 · Published November 22, 2024
Frontier AI Risk & RSI

Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations for AI R&D capabilities, and none that are highly realistic and have a direct comparison to human performance. We introduce RE-Bench (Research Engineering Benchmark, V1), which consists of 7 challenging, openended ML research engineering environments and data from 71 8-hour attempts by 61 distinct human experts. We confirm that our experts make progress in the environments given 8 hours, with 82% of expert attempts achieving a non-zero score and 24% matching or exceeding our strong reference solutions. We compare humans to several public frontier models through best-of-k with varying time budgets and agent designs, and find that the best AI agents achieve a score 4× higher than human experts when both are given a total time budget of 2 hours per environment. However, humans currently display better returns to increasing time budgets, narrowly exceeding the top AI agent scores given an 8-hour budget, and achieving 2× the score of the top AI agent when both are given 32 total hours (across different attempts). Qualitatively, we find that modern AI agents possess significant expertise in many ML topics—e.g. an agent wrote a faster custom Triton kernel than any of our human experts’—and can generate and test solutions over ten times faster than humans, at much lower cost. We open-source the evaluation environments, human expert data, analysis code and agent trajectories to facilitate future research.1

Introduction. Large language models are increasingly capable of complex programming tasks and are already being used to accelerate AI research, from generating training data to serving as programming tools [1, 2, 3, 4, 5, 6]. As these capabilities grow, there is increasing concern about the potential for AI systems to fully automate frontier AI research and development (henceforth AI R&D) with minimal human involvement [7, 8, 9].2 Multiple AI developers and governments have identified the need for evaluations that can provide advance warning of these capabilities, to allow for the implementation of key security and deployment mitigations [10, 11, 12, 13, 14].

In this work, we present RE-Bench, which aims to evaluate whether AI agents can fully automate the work of expert AI R&D researchers, using direct performance comparisons between AI agents and human experts under equivalent conditions and resources. These provide a clear logical connection to automation risk: if AI agents perform significantly worse than human ML experts given equivalent resources and conditions, then the AI agents likely cannot automate these experts’ research work.

All of the AI agents we evaluated outperform humans with a 2-hour time budget (Figure 2). However, humans display better returns to increasing time budgets, exceeding the best agent scores when given 8 hours, and continuing to improve rapidly under longer total time budgets when evaluated through best-of-k (i.e., using the result from the most successful of k independent experts) with more samples.

Qualitatively, we find that most agent solutions score close to 0 (i.e. do not improve on the reference solution), but that agents submit new solutions over 10 times faster than human experts, and occasionally find very successful approaches. Notably, both o1-preview and Claude 3.5 Sonnet find distinct solutions to our kernel optimization problem that beat the efforts of all 9 human experts (see Figure 18).

The remainder of this paper details our methodology, results and analysis. Section 2 provides background on AI R&D automation risk and current approaches to evaluation; Section 3 introduces our evaluation and the design principles behind it; Sections 4 and 5 presents our evaluation results with current frontier models; and Section 6 concludes the paper with a discussion of limitations and implications.

Related work. Effective early warning evaluations need to have a low risk of underestimating AI capabilities, while being practical and avoiding early saturation. Some key challenges of achieving this include:

• Feasibility: Achieving a high score on the evaluation has to be feasible without demanding extreme levels of capability or guesswork. To be solvable, an evaluation needs to (among other things) provide unambiguous instructions, avoid blocking technical issues, and provide enough time and resources.3 This has come up as a frequent issue with many complex evaluations; for instance, many SWE-bench [22] problems turned out to be impossible and under-specified on closer inspection. This continued to hold true even after an attempt was made to identify a verified-to-be-feasible subset [23].

• Ecological validity: To be able to calibrate risk thresholds without adding unnecessary margin or taking on excessive risk, there needs to be a clear link between evaluation results and risk levels. This is often challenging. For example, it is unclear how scores in MLEbench translate into the ability of AI agents to automate AI R&D.

Beyond benchmarking, much prior research has studied the usage of LLMs to assist with diverse aspects of machine learning research, from ideation to experimentation. LLMs are used to generate synthetic data for pretraining LLMs more efficiently and are used as reward models for fine-tuning LLMs with reinforcement learning [3, 4, 34, 35, 36].

Most relevant to our work are other evaluations and benchmarks attempting to capture challenging engineering, autonomy or research skills in LLMs, many of which have been used as a part of early warning evaluations. Table 1 contains our evaluation of several benchmarks. We find that few benchmarks provide direct human comparisons and the ones that do are often over very short time horizons in much simpler environments or QA datasets.

Of particular relevance for our work is the usage of LLMs for code generation in the domain of machine learning, either as an LLM agent or, more commonly, in a fixed scaffold that does not allow the LLM to choose what tools to use. Focusing in on code generation, LLMs have been used to do autonomous machine learning research, neural architecture search, data science problems, paper reproduction, writing research papers, and reward function design, including preference optimization for LLM fine-tuning [37, 38, 39, 40, 41, 42, 43, 44].

LLMs for science. Besides machine learning, LLMs have also been applied to R&D in other scientific domains to assist with coding, chemical synthesis, biology, and virtual scientific discovery [26, 45, 46, 47, 32].

Method. To meet the three challenges to robust and effective early warning evaluations identified in Section 2.2—feasibility, ecological validity, and resistance to saturation—we set out to create a benchmark that directly compared human experts to AI agents in realistic AI R&D environments. To enable us to iterate quickly and gather data at a reasonable cost, we decided on two practical constraints: we evaluate human experts over 8 hours, and make sure all environments can be run with 8 or fewer H100 GPUs. The environments were designed with two key pillars in mind: (i) maximize coverage of key frontier AI R&D challenges, while (ii) ensuring that human experts under the same conditions as the agents can reliably make steady progress on the task, without running into issues or hitting a score ceiling.

Covering a large variety of the challenges involved in AI R&D reduces the risk of early saturation if AI agents become capable of some but not all of the key sub-skills needed to automate AI R&D. Similarly, ensuring that environments have a high ceiling, beyond what humans can achieve within the time budget, reduces the risk that the evaluation saturates due to both agents and humans reaching The current RE-Bench suite includes seven hand-crafted novel evaluation environments (see Table 2). Each environment presents a unique ML optimization problem, where achieving a high score generally requires significant experimentation, implementation, and efficient use of compute resources. We designed these environments to have a high ceiling, but still allow significant progress with just 8 hours of time and limited hardware resources (at most 6 H100s).

• A scoring function, which defines the goal of the environment, and can be run by the agent at any time. Each time the scoring function is run a timestamped entry is added to a score log, allowing us to reconstruct progress over time. In almost all environments, the agent can see its score log, and inspect the details of the scoring function, to help it understand its goal.7 • A (simple but poorly performing) starting solution which is provided to the agent and demonstrates what a valid solution looks like. This helps clarify the environment setup and allows the agent to get to the challenging research components of the evaluation faster. For example, in ‘Optimize a Kernel,’ the agent is provided with a simple but slow Python solution.

• A reference solution, created by the task author, which scores highly. This solution is not provided to the agent, but is used in normalizing scores, and serves as an example of a good solution for researchers.

• Specification: Environment ideas were generated through discussions with ML practitioners, reading relevant literature or from past work experiences of the METR staff. To assess the idea, written descriptions of the proposed environments, scoring functions, and starting solutions were evaluated and reworked until we were optimistic that implementation would not be too challenging, and that the specification could do well on our two design pillars.

Achieving good coverage of the many challenges involved in advancing AI R&D while ensuring experts can make significant progress in just 8 hours is challenging, since real AI R&D research projects usually take months to complete. This creates a trade-off between different types of difficulty (e.g. engineering complexity, novelty, slow feedback loops), where human experts or AI agents have to split their time between different kinds of difficulty. For instance, an environment that has significant engineering complexity may not leave as much time to come up with and test novel ideas, and thus will not emphasize those skills as much.

Therefore, we have focused on creating a highly varied suite, covering different pipeline stages (pretraining, post-training, scaffolding), and types of difficulty. Table 3 highlights this variety along three key dimensions of difficulty: (i) the length of feedback loops in the environment, which determines how much the agent can rely on search/trial-and-error, (ii) the number of lines of code the agent needs to interact with and (iii) the novelty of strong solutions.

Discussion. We expect the human–AI gap in real-world AI R&D to be much larger than the gap observed on these evaluations, and find it fairly plausible that the first agents that match top human performance in these environments may still be far from capable of AI R&D automation. Specifically, we believe there Limited scope of tasks. The scale and complexity of real world AI R&D is at least 2 orders of magnitude (OOM) larger across many dimensions than those encountered in our environments (Table 7). We have seen that the human–AI gap increases over longer time horizons and qualitatively found indications that the same is true of high engineering complexity.

Cleanly defined, non-interacting tasks. Our tasks are well defined (in that each involves writing a modest amount of code to optimize a specified objective) and do not involve coordination between many interacting workstreams. We have not been able to investigate how agents deal with orchestrating parallel projects or with resolving ambiguous instructions, but given their difficulty identifying and fixing simple issues with their own environment, they may struggle with the additional communication bottlenecks, constraints and coordination difficulties required.

At the same time, it’s possible that AI agents that end up automating substantial fractions of the AI R&D process may nonetheless not exceed or even match human expert performance.

Agents are substantially cheaper than human experts: On average, our agents use ~29M input tokens and ~499K output tokens in each 8-hour run, at a cost of approximately $123. This is only a small fraction of the approximately $1,855 that we paid our human experts on average (and what a skilled human researcher costs in general), and the cost of agent runs could potentially be much lower if proper prompt caching was used.

Even if an AI agent takes longer time (or many more attempts) than a human expert to accomplish AI R&D tasks, they might nonetheless end up economically competitive with human researchers due to lower costs.

Different workflows or tooling. The tasks in RE-Bench are designed using the AI R&D workflows of human researchers. It may be possible that AI agents can automate AI research using alternative workflows that do not require performing similar tasks; for example, AI agents are capable of

Conclusion. In this work, we presented RE-Bench, a suite of environments that measure the ability of AI agents to automate AI R&D tasks. From our human expert baselines, we believe that achieving a good score on RE-Bench is feasible—though there is significant variance in the results. We also find that the environments are challenging, and most are not saturated even by the top human experts in 8 hours. We hope that these properties will allow for useful direct comparisons between human expert performance and AI agent performance on (short and self-contained) AI R&D activities.

Evaluating some current frontier AI agents we find that, when evaluated through best-of-k with 8 hours of total compute budget, they achieve scores close to the average human expert, demonstrating very impressive capabilities. However a significant gap remains compared to the top human performance in most environments, as seen in Figure 9. Monitoring whether and how quickly AI agents are bridging this gap may help predict the emergence of autonomous AI R&D automation.

Limitations. Evaluation design criteria lead to unrepresentative environments: In order to create highreliability evaluations matching our design criteria we have tried to ensure that instructions and scoring are easily understood, that significant progress is possible within 8 hours, and that all necessary resources are provided. We have also had to select for environments that are easy for us to construct and assess. These constraints make the evaluation environments less representative of real research, where unclear goals, poor instructions, slow feedback and impossible problems are common.

Lack of scaffold iteration: Different agent scaffolds or prompts might be able to achieve a much better score on our benchmark in a similar time. In particular, we expect better performance from giving the agents better tools for managing GPU resources, and from approaches that make use of a larger number of tokens, e.g., by exploring many solutions in parallel (though limited compute resources pose limitations for this approach).

Limitations in covering frontier research: Due to the limited hardware access, and because frontier AI research is increasingly siloed within large AI developers, there may be significant differences between the kinds of research these evaluations cover, and the kinds driving frontier AI progress.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do autonomous agents misreport success on failed actions? How do real-world evaluations reveal AI capabilities that benchmarks hide? Does AI-assisted research sacrifice exploration breadth for productivity gains? What human oversight must AI research systems have? Do single-axis benchmarks accurately measure agent capability for real deployment? How does AI adoption reshape collaboration patterns in knowledge work? Does AI deployment reduce or exacerbate workplace inequality and income instability? Do AI coding tools measurably improve developer productivity and code quality? How should humans and AI agents share control and decision-making? Can AI research automation sustain progress through accelerating feedback loops?