Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
Abstract. Recently, Sakana.ai introduced the AI Scientist, a system claiming to automate the entire research lifecycle and conduct research autonomously, a concept we term Artificial Research Intelligence (ARI). Achieving ARI would be a major milestone toward Artificial General Intelligence (AGI) and a prerequisite to achieving Super Intelligence. The AI Scientist received much attention in the academic and broader AI community. A thorough evaluation of the AI Scientist, however, had not yet been conducted.
We evaluated the AI Scientist and found several critical shortcomings. The system’s literature review process is inadequate, relying on simplistic keyword searches rather than profound synthesis, which leads to poor novelty assessments. In our experiments, several generated research ideas were incorrectly classified as novel, including well-established concepts such as micro-batching for stochastic gradient descent (SGD). The AI Scientist also lacks robustness in experiment execution—five out of twelve proposed experiments (42%) failed due to coding errors, and those that did run often produced logically flawed or misleading results. In one case, an experiment designed to optimize energy efficiency reported improvements in accuracy while consuming more computational resources, contradicting its stated goal. Furthermore, the system modifies experimental code minimally, with each iteration adding only 8% more characters on average, suggesting limited adaptability. The generated manuscripts were poorly substantiated, with a median of just five citations per paper—most of which were outdated (only five out of 34 citations were from 2020 or later). Structural errors were frequent, including missing figures, repeated sections, and placeholder text such as “Conclusions Here”. Hallucinated numerical results were contained in several manuscripts, undermining the reliability of its outputs.
Despite its limitations, the AI Scientist represents a significant leap forward in research automation. It produces complete research manuscripts with minimal human intervention, challenging conventional expectations of AI-generated scientific work.
Introduction. In autumn 2024, Sakana.ai, a Tokyo-based start-up that has raised $200 million in funding1, announced the “AI Scientist”.
With the AI Scientist, Sakana boldly promised the "beginning of a new era in scientific discovery" [12]. The open-source2 AI Scientist is supposed to “automate the entire research lifecycle”; i.e. it generates research ideas, designs and conducts experiments, analyzes results, writes research papers, and finally even reviews them — essentially automating the daily work of millions of researchers worldwide. According to Sakana, the AI Scientist produces research papers for approximately $15 each. Sakana acknowledges “occasional flaws” and explains further limitations in a pre-print manuscript [8]. Yet, based on Sakana’s released information, most readers will understand that the AI Scientist appears to be a fully-functional system. Especially considering Sakana’s claims that “It is worth noting that the [AI Scientist] can autonomously run the entire life cycle of machine learning research without any human intervention except for initial preparation” [13] and that the peer review system works “with near-human accuracy” [12].
Sakana’s AI Scientist is not the only AI tool that aims to support or even replace significant aspects of scientific work. A rapidly expanding body of research explores AI’s role in scientific discovery, with recent preprints examining its impact on literature retrieval, idea generation, and automated experimental design. Several studies indicate that large language models (LLMs) can already generate research ideas comparable to those of human scientists [15, 18, 19], with some findings suggesting AI may even surpass human creativity [11, 14]. Google has framed AI as ushering in a “new era of scientific discovery” [10], while specialized workshops such as “Towards Agentic AI for Science” at ICLR 2025 highlight the field’s growing momentum3.
The AI Scientist is a machine learning driven system, making it highly relevant to the machine learning research community. However, its impact extends beyond machine learning, with profound implications for the Information Retrieval (IR) community, which motivates us to write this paper. While a recent SIGIR perspective paper [20] provided valuable insights into how LLMs are transforming traditional IR tasks, we must also consider a broader question: What happens when LLMs move beyond assisting IR research and begin to conduct it autonomously? This possibility is no longer theoretical — AI systems are increasingly capable of generating research ideas, designing experiments, and even evaluating scientific work.
In this paper, we take the next step in this discussion by providing an independent, systematic assessment of the AI Scientist’s capabilities, limitations, and trajectory. Our goal is not only to evaluate its functionality but to highlight its potential role in reshaping scientific inquiry, particularly within information retrieval. By analyzing its performance across idea generation, experimentation, and manuscript production and review, we offer a comprehensive perspective on what AI-driven research tools can — and cannot — achieve today and may achieve in the future.
Our findings reveal a system that, while far from replacing human researchers, demonstrates the potential to automate significant parts of the research process. The AI Scientist does not yet fulfill its promises, struggling with methodological soundness, experimental execution, and literature retrieval. However, these limitations are technical hurdles rather than fundamental barriers — challenges that will likely be addressed as AI-driven research systems evolve.
For the IR community, the emergence of such tools presents both an opportunity and a challenge. On one hand, they offer new ways to automate literature retrieval, citation analysis, and experimental design, potentially advancing the field. On the other, they raise critical questions about the role of human researchers, the nature of scientific discovery, and the reliability and generalizability of AI-generated knowledge.
Method. 2.2 Preparing for the experiments We initially assumed the AI Scientist could autonomously conduct research based solely on a prompt. However, it requires a user-defined “template,” which significantly limits the autonomy of the AI Scientist. Such a template consist of several elements, described in the following.
- Goal and Research Direction. The research goal is specified via a prompt file (prompts.json). For our experiments, we focused on Green Recommender Systems [3], a research topic which aims to reduce the carbon footprint of recommender systems and its associated machine learning algorithms. We gave the prompt as: 14:
Listing 1. The general task provided as a JSON prompt 11https://www.reddit.com/r/learnmachinelearning/comments/1fuw2yb/has_anybody_gotten_sakana_ai_scientist_to_work/ 12HUAWEI MateBook D 16, Windows 11, Intel Core i5-12450H with integrated graphics, 16GB RAM (3733 MT/s) 13Single node with 2 AMD EPYC 7452 CPUs (32 cores, 2.35-3.35 GHz, 128 MB cache) and 256 GB DDR4 RAM (3200 MHz), running Rocky Linux 8.8 14All prompts are shortened and rephrased for brevity; the originals are available in our GitHub repository.
Manuscript submitted to ACM Evaluation of Sakana’s AI Scientist 5 1 { 2 "task_description": "The attached code trains and evaluates the FunkSVD algorithm with stochastic ↩→gradient descent for recommender systems. Your goal is to find novel ways to optimize energy- ↩→efficiency. " 3 } 2. Experimental Pipeline. The AI Scientist requires a Python-based experimental pipeline, consisting of an experiment.py, which defines dataset usage and model training as well as a plot.py, which handles result visualization. We implemented FunkSVD for collaborative filtering, training and evaluating it on an 80/10/10 split on MovieLens-100k (RMSE metric), with results visualized via line charts. This minimalist setup runs on a CPU, but larger datasets and complex models would likely increase compute costs. It is important to note that the AI Scientist requires this pipeline in a special format. Users cannot simply input any Python code but would have to adjust their existing code to be able to be used by the AI Scientist.
- Seed Ideas. The AI Scientist further requires user-provided seed ideas. We supplied two:
(1) Adaptive Learning Rates for SGD. Aims to accelerate convergence, reducing computational cost and energy use. Though not novel, this ideas demonstrates our intended research direction.
(2) E-Fold Cross-Validation. An alternative to k-fold cross-validation that dynamically selects an optimal eto balance computational efficiency and reliability. This idea is also not entirely novel, but was proposed only recently[1, 4, 9].
The seed ideas must include values for “interestingness,” “feasibility,” and “novelty.” However, we observed that these values had no apparent impact on the AI Scientist’s processing.
Listing 2. e-fold Cross Validation; one of the two seed ideas 1 { 2 "Name": "e_fold_cv", 3 "Title": "E-Fold Cross-Validation for Model Performance Estimation", 4 "Experiment": "Develop an alternative to k-fold cross-validation named e-fold cross-validation. Instead ↩→ of a static k it uses an intelligently chosen or dynamically adjusted paramter e to optimize ↩→the number of folds and to minimize computational energy while maintaining reliable model ↩→performance estimates.", 5 "Interestingness": 8, 6 "Feasibility": 7, 7 "Novelty": 7 8 } 2.3 Idea Generation Once provided with required input, the AI Scientist generated the following ten research ideas:
Manuscript submitted to ACM 6 Beel & Kan et al.
- Factor pruning for energy-efficient matrix factorization 2. Dynamic factor adjustment in matrix factorization 3. Green metrics for sustainable recommender systems 4. Quantization techniques for energy-efficient SGD 5. Time-aware early stopping 6. Interaction-priority SGD for energy efficiency 7. Micro-batching for energy-efficient SGD 8. Sparse-aware SGD for matrix factorization 9. Cluster-aware SGD for matrix factorization 10.
Discussion. A key takeaway from our study is that we believe AI systems like the AI Scientist will meaningfully contribute to the scientific critique process in the near future, benefiting both authors and reviewers. Researchers could leverage AI-generated feedback to iteratively refine their work through adversarial loops in which an AI challenges their hypotheses, methodologies, and analyses. Similarly, peer reviewers could use AI-assisted prompts to enhance their critiques, ensuring consistency and depth. Some academic conferences, such as ICLR, have already incorporated AI- generated review suggestions into their workflows. However, our findings indicate that AI-generated reviews by the AI Scientist often focus on surface-level critiques, while failing to detect deeper methodological flaws. This may make the job of second-level reviewers (such as area chairs or meta-reviewers) more important, as AI-generated reviews may lack depth but be well-formatted and directly address the review form’s criteria. Such consolidators will have to look beyond surface-level analyses to judge whether a work holds promise.
Another promising role for the AI Scientist lies in replication and validation. Reproducibility is a cornerstone of scientific credibility, as highlighted by recent initiatives in recommendation systems (e.g., the Best Paper Award at RecSys) and NLP (e.g., ReproHum initiatives). AI could be employed to establish proof of work, ensuring that findings are replicable and that experimental procedures are verifiable. To achieve this, AI-driven replication efforts must prioritize transparency through open-weight and open-data models. Additionally, AI can support automated metadata creation and tagging (e.g., Dublin Core, Open Data Initiative) and facilitate structured data deposits in version-controlled repositories, providing verifiable provenance for scientific claims. These mechanisms incentivize transparent and reproducible research, potentially leading to new prestige metrics or blockchain-based certification systems for replication integrity.
Despite these potential benefits, integrating AI into scientific research presents ethical and practical challenges.
AI models inherit biases from historical data and cannot independently distinguish between scientific quality and consensus. This limitation raises concerns that AI tools might reinforce outdated methodologies or amplify biases in peer review and publishing. Additionally, junior researchers may develop an overreliance on AI-generated suggestions, leading to automation bias that stifles innovation and independent thought. Research further indicates that AI can enhance scientific productivity, as shown in [17], yet its impact on researcher satisfaction remains complex. In one Manuscript submitted to ACM Evaluation of Sakana’s AI Scientist 15 study, while AI-assisted workflows increased efficiency, 82% of scientists reported lower job satisfaction afterwards.
This paradox exemplifies a broader dilemma—if AI can outperform humans in key research tasks, what remains for human scientists to do? Addressing these risks requires AI-assisted workflows that actively encourage users to critically reassess assumptions and explore alternative perspectives.
One pressing concern is the potential for mass AI-generated paper submissions. Tools like the AI Scientist could enable researchers to generate large volumes of seemingly novel but low-value papers by tweaking existing codebases and experimental parameters. Such submissions could overwhelm academic venues, burdening already-overworked reviewers. The precedent set by AI-generated nonsense papers being accepted into conferences, such as those produced by SciGen, demonstrates the urgency of this issue.
Conclusion. In summary, while the AI Scientist falls short of its grand claims, it offers a glimpse into the future of AI-driven scientific discovery and towards ‘Artificial Research Intelligence’ (ARI). With responsible development and integration, AI could enhance knowledge generation, improve reproducibility, and streamline the research process in ways we are only beginning to explore. However, realizing this potential requires careful consideration of AI’s limitations, ethical challenges, and long-term implications for the scientific community. The time to act and to participate in evaluating, exploring, discussing and developing AI-driven research agents is now!
Limitations. 2.8 Limitations Our own evaluation had limitations: we used a single dataset (MovieLens-100k) and one research domain (Green Recommender Systems), relied on fixed seed ideas and one template. While these constraints ensured a controlled assessment, they may affect generalizability and reproducibility. Despite this, we believe our findings remain meaningful.
Manuscript submitted to ACM
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI systems discover fundamental improvements to their own architectures? How do hallucinated citations emerge in AI scholarly output? Can AI systems perform peer review as effectively as humans?- Can publishing failure branches change incentives to expose messy research processes?
- How can automated review scale with the flood of AI-generated papers?
- What accountability structures should replace detection when AI automation increases in peer review?
- How do closed-loop automated venues differ from human-in-the-loop review taxonomies?
- What collaboration model between humans and AI best serves peer review?
- How much has peer review workload grown at major conferences?
- Can automated AI systems assess novelty as well as human reviewers?
- Should AI research papers require dedicated automated review systems instead of existing journals?
- Could AI improve peer review rigor and catch human-missed errors?
- Could AI feedback work as a substitute for human peer review entirely?
- What effects do preprint servers have on scientific consensus formation?
- How can arXiv and journals scale quality control for AI-generated research?
- Do AI-generated research reviews score papers higher than human reviewers do?
- Does AI content in reviews correlate with differences in paper quality control?
- Why should AI research prompts be subject to peer review before use?
- What role should humans play in reviewing and approving AI-generated research?