Towards End-to-End Automation of AI Research
The automation of science is a long-standing ambition in the field of AI (Buchanan and Feigenbaum, 1978; Lenat, 1977). While the community has made significant progress in automating individual components of the scientific process, a system that autonomously navigates the entire research lifecycle—from conception to publication—has remained out of reach. Here, we present the strongest demonstration to date toward automating the entire process end-to-end. We present The AI Scientist, which creates research ideas, writes code, runs experiments, plots and analyzes data, writes the entire scientific manuscript and performs its own peer review. Its ideas, execution, and presentation are of sufficient quality to produce a manuscript generated by an AI system that passes the first round of peer review at a major machine learning conference workshop. The workshop has an acceptance rate of 70 percent. Our system leverages modern foundation models (Anthropic, 2024; Llama Team, 2024; OpenAI, 2023) within a complex agentic system. We evaluate The AI Scientist in two settings: a focused mode using human-provided code templates as an initial scaffold to conduct research on a specific topic, and a template-free, open-ended mode that leverages agentic search for wider scientific exploration (Chan et al., 2025; Jiang et al., 2025). Both settings produce diverse ideas and automatically test, report on, and evaluate them. This achievement demonstrates AI’s growing capacity for scientific contribution and signifies a potential paradigm shift in how research is conducted. As with any impactful new technology, there could be significant risks, including taxing overwhelmed review systems and adding noise to scientific literature. However, if developed responsibly, such autonomous systems could greatly accelerate scientific discovery.
Introduction. AI has long been used to aid scientific discovery, an ambition with deep roots in the history of the field (Langley, 2024; Langley et al., 1987; Lenat, 1977; Lenat and Brown, 1984; Waltz and Buchanan, 2009). Prior to the rise of large language models (LLMs), AI was limited to helping with specific, narrow tasks, such as discovering chemical structures (Buchanan and Feigenbaum, 1978), finding mathematical proofs (Lenat, 1977), discovering novel materials (Merchant et al., 2023; Pyzer-Knapp et al., 2022; Szymanski et al., 2023), and predicting the 3D shape of proteins (Hayes et al., 2025; Jumper et al., 2021). Other systems focused on analyzing pre-collected datasets to find novel insights (Falkenhainer and Michalski, 1986; Ifargan et al., 2025; Langley et al., 1987). However, with the recent advent of powerful and general foundation models, AI’s role has expanded to assist with a wider array of research activities. For example, LLMs now help with generating novel hypotheses (Faldor et al., 2024; Girotra et al., 2023; Hu et al., 2025; Lehman et al., 2023; Lu et al., 2025), writing literature reviews (Baek et al., 2025; Wang et al., 2024b), and coding experiments (Huang et al., 2024; Lu et al., 2024a; Ma et al., 2023; Zhang et al., 2025). Despite these advances in automating individual components, a system that autonomously navigates the entire research lifecycle—from conception to publication—has remained out of reach until now.
Method. The AI Scientist sequentially completes four main phases (Figure 1A): In the first phase, The AI Scientist is prompted to iteratively grow an archive (Mouret and Clune, 2015) of high-level research directions and hypotheses it can explore within a user-specified machine learning research subfield (an example progression is visualized in Supplementary Section C.4). For each direction, it generates a descriptive title, explanation of its reasoning for what the idea is and why it’s interesting to pursue it, and a proposed experimental plan (Supplementary Sections A.1.1 and A.2.6). After idea generation, The AI Scientist filters ideas by connecting the language model to the Semantic Scholar API (Fricke, 2018) and web access as tools (Schick et al., 2024). This allows The AI Scientist to discard any idea that is too similar to existing literature. The second phase of The AI Scientist executes the proposed experiments and then visualizes their results for the downstream write-up. We tested two different variants of experiment execution: (1) Template-Based: The AI Scientist is provided with a starting code template that reproduces a training run from a popular algorithm. The AI Scientist then executes the proposed experiment plan in linear order (Supplementary Section A.1). (2) Template-Free: Alternatively, The AI Scientist can generate an initial starting code script by itself. In this case, experimentation includes additional stages for optimizing the code it writes from scratch, and experiment execution leverages additional test-time compute with tree search (see methods). After each experiment, The AI Scientist is given the results and is prompted to take notes in the style of an experimental journal for future planning and writeup. The third phase of The AI Scientist produces a concise write-up of its research in the style of a standard machine learning conference paper. The AI Scientist is prompted to fill in a blank LaTeX conference template section by section, using its notes and plots (see methods). To construct the related work section and add citations throughout the manuscript, the system queries the Semantic Scholar (Fricke, 2018) API for relevant literature, comparing its findings against the generated manuscript over 20 rounds. For each potential citation, the system generates a textual justification for its inclusion, which informs The AI Scientist on how to use the reference appropriately within the manuscript. Finally, the paper generated by The AI Scientist undergoes a review by the Automated Reviewer to automatically evaluate the scientific quality of the conducted research.
The Automated Reviewer provides reviews based on the top-tier Neural Information Processing Systems (NeurIPS) conference review guidelines (NeurIPS Program Chairs, 2022). The output contains numerical scores (soundness, presentation, contribution, overall, and reviewer confidence), lists of weaknesses and strengths, as well as a binary decision (accept or reject). The Automated Reviewer’s pipeline consists of an ensemble of five reviews, followed by a meta-review where the model acts as an Area Chair to make a final decision conditioned on all five reviews (Supplementary Section A.3). We compared Automated Reviewer decisions with ground truth data for ICLR papers, extracted from the publicly available OpenReview dataset (González-Márquez and Kobak, 2024). As shown in Table 1, the Automated Reviewer’s agreement with human paper assessments is comparable to inter-human agreement measured by F1 and balanced accuracy as reported in the NeurIPS 2021 consistency study (Beygelzimer et al., 2021), which measured agreement between human reviewers on a comparable set of submissions (Supplementary Section A.3). This demonstrates its ability to replicate the collective judgment of human reviewers with high fidelity. These results are statistically significant (non-parametric bootstrap test (Efron and Tibshirani, 1993) and two-sample z-test (Lehmann, 1959); Supplementary Section A.3). Next, to investigate the effect of potential
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI systems perform peer review as effectively as humans?- Should AI research papers require dedicated automated review systems instead of existing journals?
- Can workshop acceptance rates reliably measure AI research quality compared to main conferences?
- Could AI feedback work as a substitute for human peer review entirely?
- What role should humans play in reviewing and approving AI-generated research?
- Do humans or AI perform better at different research stages?
- How does data availability shape which scientific questions AI systems tackle?
- Where does AI assistance become reliable versus prone to failure in science?
- How do template requirements limit AI research systems from true autonomy?
- What role should human experts play in AI-driven research ideation loops?
- What citation mistakes appear in fully autonomous AI research pipelines?
- How do researchers currently check whether an autonomous system's novelty claims are actually valid?
- What would a practical reviewer checklist for autonomous research systems need to include?
- Can AI agents themselves become reliable reviewers of other autonomous research systems?
- Can AI agents align their ideas with future research directions as well as humans do?
- Can human-AI collaboration preserve scientific breadth while improving individual productivity?
- How do autonomous science systems preserve competing hypotheses without a central planner?
- How do multi-agent writing systems maintain consistency across scientific manuscript sections?
- Does AI research acceleration compound into faster field-wide progress over time?
- Can recursive feedback loops turn AI research automation into genuine progress?
- What domains allow autonomous AI discovery because verification is fast enough?
- How does feedback latency from physical experiments shape AI system autonomy in research?