Kosmos: An AI Scientist for Autonomous Discovery
Abstract Data-driven scientific discovery requires iterative cycles of literature search, hypothesis generation, and data analysis. Substantial progress has been made towards AI agents that can automate scientific research, but all such agents remain limited in the number of actions they can take before losing coherence, thus limiting the depth of their findings. Here we present Kosmos, an AI scientist that automates data-driven discovery. Given an open-ended objective and a dataset, Kosmos runs for up to 12 hours performing cycles of parallel data analysis, literature search, and hypothesis generation before synthesizing discoveries into scientific reports. Unlike prior systems, Kosmos uses a structured world model to share information between a data analysis agent and a literature search agent. The world model enables Kosmos to coherently pursue the specified objective over 200 agent rollouts, collectively executing an average of 42,000 lines of code and reading 1,500 papers per run. Kosmos cites all statements in its reports with code or primary literature, ensuring its reasoning is traceable. Independent scientists found 79.4% of statements in Kosmos reports to be accurate, and collaborators reported that a single 20-cycle Kosmos run performed the equivalent of 6 months of their own research time on average. Furthermore, collaborators reported that the number of valuable scientific findings generated scales linearly with Kosmos cycles (tested up to 20 cycles). We highlight seven discoveries made by Kosmos that span metabolomics, materials science, neuroscience, and statistical genetics. Three discoveries independently reproduce findings from preprinted or unpublished manuscripts that were not accessed by Kosmos at runtime, while four make novel contributions to the scientific literature.
Introduction. Data-driven discovery consists of iterative steps of literature search, hypothesis generation, and data analysis to draw new scientific conclusions from datasets. Due to their proficiency in programming and interdisciplinary reasoning, large language model (LLM) agents have the potential to automate data-driven discovery across domains.
Robin [1], a system we previously reported, performs automated cycles of literature search and data analysis to propose evidence-based hypotheses, but has limited context sharing between its agents and is primarily tailored for therapeutics development. Sakana’s AI Scientist [2] autonomously forms hypotheses, iteratively conducts computational experiments, and writes and reviews manuscripts about its results, but remains limited to machine learning research. Google’s AI co-scientist [3] conducts iterative cycles of reasoning to generate scientific hypotheses, but does not perform or analyze experiments. The Virtual Lab [4] successfully designed novel nanobodies that neutralize SARS-CoV-2, and may be extensible to other domains, but lacks exploratory data analysis capabilities.
Here, we present Kosmos, an AI scientist that automates data-driven discovery across a wide range of scientific disciplines. Given an open-ended objective and a dataset, Kosmos performs iterative cycles of parallel data analysis, literature search, and hypothesis generation, and summarizes its discoveries in scientific reports (Figure 1a). At each cycle, Kosmos launches several parallel instances of two general-purpose Edison Scientific agents, a data analysis agent [5] and a literature search agent [6], with each instance assigned to a specific task that is aligned with the end objective. Kosmos shares and synthesizes information among these agents by continuously updating a structured world model, which enables Kosmos to execute an average of 42,000 lines of code across 166 data analysis agent rollouts and read 1,500 full-length scientific papers across 36 literature review agent rollouts per run (Figure 1b). This is a 9.8x increase in code generation compared to Robin. Consolidating information in the world model further allows every claim in a Kosmos scientific report to be directly linked to the data analysis or source from which it originated, ensuring that Kosmos’ reasoning is traceable.
We report seven discoveries made by Kosmos: three discoveries made by Kosmos reproduce findings from preprinted or unpublished manuscripts, while the remaining four make novel contributions to the scientific literature. Each discovery is derived from a unique data type and field and is corroborated by independent analysis from a domain expert. Among the examples below, Kosmos identified a novel, clinically relevant mechanism of neuronal aging, and generated novel statistical evidence that high circulating levels of superoxide dismutase 2 (SOD2) may causally reduce myocardial fibrosis in humans. Together, these discoveries illustrate a system that can autonomously reproduce, refine, and generate data-driven discoveries with the rigor and transparency essential for advancing scientific understanding.
Method. The core advancement in Kosmos is the use of a structured world model to manage the output of a large number of agents running in parallel. Kosmos is initiated with a research objective and a dataset, which are specified by a scientist. Kosmos attempts to complete the research objective by using LLMs, data analysis agents, literature search agents, and the world model to perform iterative discovery cycles. In each cycle, Kosmos executes up to ten literature search and analysis tasks, and subsequently updates the world model with summaries of the task outputs. Kosmos then queries the world model to propose literature search and data analysis tasks to be completed in the next cycle. This context management strategy allows Kosmos to explore many different research avenues simultaneously, and run for eight times as many iterations than existing systems [1, 2, 7]. Once Kosmos believes it has completed the research objective, it synthesizes key discoveries into three or four scientific reports. Each statement and figure in the report cites either a publication found by the literature search agent or a Jupyter notebook created by the data analysis agent.
Discussion. To our knowledge, Kosmos is the first example of an AI Scientist that can carry out months of work in a single run, combining closed-loop literature search, data analysis, and world model updates to autonomously make discoveries in multiple fields. Kosmos does this by using structured world models to manage context between agents, allowing it to deploy hundreds of agent rollouts, write tens of thousands of lines of code, and read thousands of papers to complete a single research objective. Kosmos performs large-scale, unbiased exploration of high-dimensional datasets to successfully reproduce known work (Figures 2-4), refine and synthesize existing knowledge (Figures 5-7), and make novel discoveries (Figure 8). Because Kosmos uses a dynamically updated world model to deploy two general-purpose scientific agents, Kosmos is able to operate in any domain. While we have demonstrated its utility in metabolomics (Figure 2), materials science (Figure 3), connectomics (Figure 4), statistical genetics (Figures 5, 6), proteomics (Figure 7), and transcriptomics (Figure 8), we anticipate that Kosmos can be applied to diverse data-rich fields. Finally, each statement in a Kosmos report is supported by a piece of code or primary literature citation. This allows the entire discovery process to be transparent and facilitates independent verification or replication of any finding, as presented for the discoveries here.
Kosmos is designed not to replace human scientists, but to augment and accelerate their work. A Kosmos-integrated pipeline begins with human-generated and -curated high-quality dataset and ends with human interpretation and critical evaluation of the results. Thus, the quality and format of input datasets have a significant impact on the output discoveries. For instance, preliminary runs for discoveries in Figures 2, 5 and 6 yielded different focuses and results depending on the type of input data pre-processing (data not shown). Throughout our testing, collaborators found Kosmos provided the best scientific insight when given clearly labeled, well-formatted, and properly normalized data.
Following a Kosmos run, the scientist’s role shifts to evaluation and interpretation of results. This human oversight is crucial, as independent evaluators noted that Kosmos tends to make excessively strong claims and can sometimes veer in unexpected trajectories (Figure 1c). Together, scientists and Kosmos form a rapid feedback loop where initial human findings feed AI-generated hypotheses which are then refined to guide subsequent searches, ensuring Kosmos’ analytical power remains directed toward accurate and meaningful scientific goals.
Conclusion. Taken together, our results demonstrate an AI scientist that can autonomously conduct extended research investigations that produce discoveries validated by domain experts across multiple scientific fields. The world model enables coordination of parallel agent trajectories at scales that would require months of human effort, while maintaining complete traceability of all scientific claims. With further training, Kosmos has the potential to significantly scale data-driven discovery.
Limitations. Kosmos has several limitations that highlight opportunities for future development. First, although 85% of statements derived from data analyses were accurate, our evaluations do not capture if the analyses Kosmos chose to execute were the ones most likely to yield novel or interesting scientific insights. Kosmos has a tendency to invent unorthodox quantitative metrics in its analyses that, while often statistically sound, can be conceptually obscure and difficult to interpret. Similarly, Kosmos was found to be only 57% accurate in statements that required interpretation of results, likely due to its propensity to conflate statistically significant results with scientifically valuable ones. Given these limitations, the central value proposition is therefore not that Kosmos is always correct, but that its extensive, unbiased exploration can reliably uncover true and interesting phenomena. We anticipate that training Kosmos may better align these elements of “scientific taste” with those of expert scientists and subsequently increase the number of valuable insights Kosmos generates in each run.
Second, identifying valuable discoveries Kosmos made is a time-intensive process that relies on human scientists with significant domain expertise. An average Kosmos report contains 3-4 discovery narratives, and each discovery narrative contains 25 claims based on 8-9 agent trajectories.
Lines of inquiry this paper opens 12
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does AI-assisted research sacrifice exploration breadth for productivity gains?- Can explicit collaboration rules in hypothesis generation be tested and varied independently?
- How do autonomous science systems preserve competing hypotheses without a central planner?
- Does decentralized coordination preserve more research hypotheses than a central world model planner?
- Do research agents mostly reproduce known techniques or discover novel solutions?
- Can agentic AI systems handle judgment-intensive tasks in science?
- What role should human experts play in AI-driven research ideation loops?
- How should AI tools integrate into wet-lab biology discovery workflows?