The AI Scientist Generates its First Peer-Reviewed Scientific Publication
Source: Sakana AI · 2025-03-12
A paper produced by The AI Scientist-v2 passed the peer-review process at a workshop in a top international AI conference.
We are proud to announce that a paper produced by The AI Scientist passed the peer-review process at a workshop in a top machine learning conference.
The paper was generated by an improved version of the original AI Scientist, called The AI Scientist-v2. This paper was submitted to an ICLR 2025 workshop that agreed to work with our team to conduct an experiment to double-blind review AI-generated manuscripts. We selected this workshop because of its broader scope, challenging researchers (and our AI Scientist) to tackle diverse research topics that address practical limitations of deep learning.
We worked with the ICLR workshop organizers, and agreed that we would submit 3 AI-generated papers into the workshop for peer-review. The reviewers were informed about the possibility and likelihood that papers they are reviewing might be AI generated (3 out of 43 papers) but not if the papers assigned to them were actually AI generated or not (for details, see the ICLR workshop’s Reviewer Guidelines).
Critically, the AI-generated papers we submitted were entirely generated end-to-end by AI, without any modifications from humans. The AI Scientist-v2 came up with the scientific hypothesis, proposed the experiments to test the hypothesis, wrote and refined the code to conduct those experiments, ran the experiments, analyzed the data, visualized the data in figures, and wrote every word of the entire scientific manuscript, from the title to the final reference, including placing figures and all formatting.
We, as the humans overseeing this research, merely gave it the broad topic to perform research on (because the topic should be relevant to the workshop we submitted to) and picked 3 AI-generated papers to submit. We chose this number following discussions with the workshop organizers to avoid overburdening reviewers.
We looked at the generated papers and submitted those we thought were the top 3 (factoring in diversity and quality—We conducted our own detailed analysis of the 3 papers, please read on in our analysis section). Of the 3 papers submitted, two papers did not meet the bar for acceptance. One paper received an average score of 6.33, ranking approximately 45% of all submissions. These scores are higher than many other accepted human-written papers at the workshop, placing the paper above the average acceptance threshold. Specifically, the scores were:
However, as we will highlight in the next section about the Importance of Transparency and Ethical Code of Conduct, it was determined ahead of time, as part of our experiment protocol, that even if papers by The AI Scientist were accepted, we would withdraw them before they were actually published. This is because they were AI-generated, and the AI and scientific communities have not yet decided whether we want to publish AI-generated manuscripts in the same venues.
For transparency, because this paper was withdrawn after the peer-review process, the ICLR workshop organizers did not perform any additional meta-review on the paper, as they were already aware of this experiment. Hence, even though the paper received an average score of 6.33, it is still possible that a meta reviewer (in this case, the workshop organizers), in theory, could have rejected this paper.
The original AI Scientist represented the first time AI generated entire scientific manuscripts. To our knowledge, this is the first time a fully AI-generated paper was good enough to pass a standard scientific peer-review process like the one described.
The AI Scientist-v2, after being given a broad topic to conduct research on, generated a paper titled “Compositional Regularization: Unexpected Obstacles in Enhancing Neural Network Generalization”. This paper reported a negative result that The AI Scientist encountered while trying to innovate on novel regularization methods for training neural networks that can improve their compositional generalization. This manuscript received an average reviewer score of 6.33 at the ICLR workshop, placing it above the average acceptance threshold.
We believe it is important for the scientific community to study the quality of AI-generated research, and one of the best ways to do so is to submit a small sample of it to the same rigorous peer-review processes we use to assess human-generated science (provided one has permission from those managing such processes).
Furthermore, our AI-generated papers will not be made accessible on OpenReview’s public forum. This is because for the purpose of this particular experiment, the ICLR conference organizers, ICLR workshop organizers and ourselves have agreed that AI-generated papers will be withdrawn from further consideration, and automatically desk-rejected after the peer-review process has been completed.
We as a community also need to develop norms regarding AI-generated science, including when and how to declare that a paper is fully or partially AI-generated, and at what point in the process. We will share more details on these issues in our forthcoming paper, but at a high level, we believe in providing as much transparency as possible regarding what is AI-generated, although there are difficult questions about whether the science should be judged on its own merits first to avoid bias against it.
Going forward, we will continue to exchange opinions with the research community on the state of this technology to ensure that it does not develop into a situation in the future where its sole purpose is to pass peer review, thereby substantially undermining the meaning of the scientific peer review process.
We note that while our AI Scientist has successfully generated peer-reviewed work, the venue in which the work is presented is at the workshop track, rather than at the main conference track. We also reiterate that only 1 out of the 3 generated papers had been accepted at this workshop.
Typically, workshop papers present preliminary findings that are less refined compared to main conference submissions, and in fact, many conference papers started off as a workshop paper. As we will describe later in our analysis section) below, we, as human AI researchers, also conducted our own internal reviews of the 3 papers, but concluded that none of them passed our internal bar for an ICLR conference track publication.
The acceptance rates at the main conference at a top machine learning conference like ICLR, ICML and NeurIPS are typically in the 20-30% range, while at the workshops like the one we submitted to, hosted along top ML conferences, have acceptance rates in the 60-70% range. In future work, we intend to improve our process to produce even higher quality scientific papers that may pass the bar of top-tier conferences.
In addition to the peer-review process, as human AI researchers, we also conducted our own analysis and reviewed all of the 3 AI-generated papers. We treated the 3 papers as if they were manuscripts submitted to the main ICLR conference track (which has a higher bar for acceptance), and our team wrote comprehensive reviews for each generated paper.
The AI Scientist occasionally made embarrassing citation errors. For instance, here, we found that it incorrectly attributed “an LSTM-based neural network” to Goodfellow (2016) rather than to the correct authors, Hochreiter and Schmidhuber (1997).
Ultimately, we concluded that none of the 3 papers passed our internal bar for what we believe would qualify as an accepted ICLR conference track paper, in their current forms. However, we believe that the papers we sent to the workshop contain interesting, original, though preliminary ideas that can be developed further, hence we believe they may qualify for the ICLR workshop track.
We believe the next generations of The AI Scientist will usher in a new era in science. That AI can generate an entire scientific paper that passes peer-review at a top-tier ML workshop conveys very promising early signs of progress. But this is just the beginning. We expect AI to continue to improve, potentially exponentially. At some point in the future, AI will probably be able to generate papers at and beyond human levels, including at the highest level of scientific publishing. We predict The AI Scientist and systems like it will create papers worthy of acceptance not only at top ML conferences, but also in the top journals in science.
Ultimately, we believe what matters most is not how AI science is judged vs. human science, but whether its discoveries aid in human flourishing, such as curing diseases or expanding our knowledge of the laws that govern our universe. We look forward to helping usher in this era of AI science contributing to the betterment of humanity.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI systems perform peer review as effectively as humans?- Did adding AI reviews actually change peer review decisions or paper outcomes?
- Can human reviewers detect when papers have been rewritten by AI?
- Do AI reviews depend more on writing style than scientific merit?
- Can workshop acceptance rates reliably measure AI research quality compared to main conferences?
- Can automated review systems catch deep methodological flaws or only surface issues?
- Could AI improve peer review rigor and catch human-missed errors?
- Could AI feedback work as a substitute for human peer review entirely?
- How can arXiv and journals scale quality control for AI-generated research?
- Can institutional statements alone correct misconceptions from unreviewed papers?
- Do AI-generated research reviews score papers higher than human reviewers do?
- Can human reviewers reliably detect AI-written peer review text by sight?
- How often do AI systems produce papers with undetected factual errors?
- How do automated reviewers detect flaws that human experts miss in manuscripts?
- How do citation errors in AI-generated papers differ from human hallucinations?
- Can novelty filters using literature search prevent AI-generated research from duplicating prior work?
- What role should humans play in reviewing and approving AI-generated research?
- How do template requirements limit AI research systems from true autonomy?
- Are paper mills using NHANES data to automate single-factor research?
- When should domain experts verify AI research claims before publication?
- What citation mistakes appear in fully autonomous AI research pipelines?