Robin: A multi-agent system for automating scientific discovery

Paper · arXiv 2505.13400 · Published May 19, 2025
Domain Specialization in LLMs

Scientific discovery is driven by the iterative process of background research, hypothesis generation, experimentation, and data analysis. Despite recent advancements in applying artificial intelligence to scientific discovery, no system has yet automated all of these stages in a single workflow. Here, we introduce Robin, the first multi-agent system capable of fully automating the key intellectual steps of the scientific process. By integrating literature search agents with data analysis agents, Robin can generate hypotheses, propose experiments, interpret experimental results, and generate updated hypotheses, achieving a semi-autonomous approach to scientific discovery. By applying this system, we were able to identify a novel treatment for dry age-related macular degeneration (dAMD), the major cause of blindness in the developed world. Robin proposed enhancing retinal pigment epithelium phagocytosis as a therapeutic strategy, and identified and validated a promising therapeutic candidate, ripasudil. Ripasudil is a clinically-used rho kinase (ROCK) inhibitor that has never previously been proposed for treating dAMD. To elucidate the mechanism of ripasudilinduced upregulation of phagocytosis, Robin then proposed and analyzed a follow-up RNA-seq experiment, which revealed upregulation of ABCA1, a critical lipid efflux pump and possible novel target. All hypotheses, experimental plans, data analyses, and data figures in the main text of this report were produced by Robin. As the first AI system to autonomously discover and validate a novel therapeutic candidate within an iterative lab-in-the-loop framework, Robin establishes a new paradigm for AI-driven scientific discovery.

Introduction. Advances in our ability to measure, perturb, and model biological systems have resulted in rapid growth of our collective scientific knowledge [1]. Yet, complementary technologies to interpret, synthesize, and generate hypotheses from this knowledge have lagged behind [2]. Artificial intelligence (AI) systems based on large language models (LLMs) show promise for automating this knowledge synthesis process and accelerating scientific discovery. As a primary goal of biomedical research is the development of new treatments for disease, our ability to produce new therapeutics may be the ultimate beneficiary of these approaches. Drug development heavily relies on a confluence of biological, clinical, and pharmaceutical expertise, and is limited by the rate at which these experts can synthesize the scientific literature [3].

The repurposing of existing drugs for new indications presents a promising application space for LLM systems. The history of drug repurposing often shows a pattern: while insights often existed in scientific literature, only after a significant lag did that knowledge crystallize into a new treatment. For example, dabrafenib, an inhibitor of BRAF kinase that is used in various cancers with mitogenic mutations in BRAF, is being repurposed to prevent hearing loss [4, 5]. While its molecular action was well characterized by 2010 [6, 7, 8], dabrafenib’s otoprotective effects were only discovered 10 years later via unbiased high-throughput screening [6]. This delayed discovery occurred despite dabrafenib’s otoprotective effects being a direct result of its known inhibition of BRAF [4, 9, 5], suggesting that more repurposing opportunities could be identified through logical connection of existing biological insights in the literature. Further cases across medicine – from ketamine (22 year lag [10, 11]) to leucovorin (5 year lag [12, 13]) to KarXT (13 year lag [14, 15]) – underscore how repurposing efforts are frequently realized years after core insights are documented.

Such delays in connecting existing insights to new therapeutic applications highlight the challenge of synthesizing disparate scientific knowledge. Trained on data across many fields, large language models (LLMs) are able to store and recall information on a wide variety of scientific topics and thus transcend the limitations of individual human knowledge. Previous work has shown that fine-tuned LLMs and specialized RAG systems can exceed human performance on retrieving and summarizing information from the scientific literature. These advances raise the possibility that LLM systems could be used for novel hypothesis generation [16, 17, 18, 19].

Here, we introduce Robin, the first multi-agent system for scientific discovery that integrates novel hypothesis generation with experimental data analysis in one continuous workflow (Figure 1A). Robin utilizes specialized language agents for literature search (Crow and Falcon) and data analysis (Finch) to enable semi-autonomous scientific discovery [27, 28]. Though this system could be applicable to scientific discovery across disciplines, in this report, we focus on its potential in therapeutics. Given a disease, Robin automatically identifies relevant in vitro assays that model key disease mechanisms and proposes specific drug candidates to evaluate in these experimental models. Researchers next conduct the experiments and provide the resulting data to Robin for autonomous analysis. Robin then interprets the results of this analysis to generate a new round of therapeutic candidates. Through this process, Robin drives an iterative therapeutics development cycle where hypotheses are generated, tested, analyzed, and refined based on experimental results. The key intellectual steps of the scientific method are thus automated while coordinating with scientists throughout the experimental loop. By connecting literature-based hypothesis generation with experimental data analysis in a continuous feedback system, Robin represents the first complete implementation of AI-driven scientific discovery.

To demonstrate Robin’s ability to generate and refine novel therapeutic hypotheses, we attempted to identify potential new treatments for dry age-related macular degeneration (dAMD). dAMD is the leading cause of irreversible sight loss in developed countries, yet limited treatment options are available. In the U.S. alone, 1.5 million people have vision-threatening dAMD, and 600,000 are legally blind due to AMD, a figure projected to almost triple by 2050 due to an aging population [29, 30]. By applying Robin to discover novel therapeutic candidates for dAMD, this work represents a first step towards AI-generated discovery in scientific research.

Related work. Several LLM systems have recently been developed to automate hypothesis generation [20, 21, 22, 23, 24, 25]. Specialized systems have also been developed to automate specific tasks in drug discovery, such as prediction of pharmacological properties and safety profiles [25, 22]. These systems have demonstrated they can generate reasonable hypotheses by utilizing multi-agent architectures that decompose scientific reasoning into discrete manageable subtasks [25], domain-specific fine-tuning [24, 26], integration of external tools [24, 25], and incorporation of human feedback [25]. However, none of these systems are currently capable of fully automating the key intellectual steps of the scientific process, including generating hypotheses and experimental strategies, analyzing results from the experiments, and refining hypotheses in light of new data.

Method. Robin integrates multiple language agents in a structured workflow to generate therapeutic candidates for a given disease (Figure 1A,B). Crow and Falcon are literature search agents based on PaperQA2 that conduct concise and deep literature summaries, respectively [27]. PaperQA2 achieves expert-level performance in information retrieval and summarization, with access to scientific literature, clinical trial reports, and the Open Targets Platform [31]. Finch is a scientific data analysis agent that performs analyses of experimental data from assays, such as RNA-seq and flow cytometry (Figure 1C) [28]. By coordinating these agents to identify novel therapeutics, Robin enables an experimentally-guided system that drives the process of scientific discovery. An example of the Robin hypothesis generation workflow is shown in Supplementary Figure S1.

Robin generates therapeutic candidates via scientific synthesis with Falcon literature search and if provided, experimental results.

Experimental Assay Selection Robin first selects an experimental assay by conducting literature review with Crow and generating assay proposals.

Experimental insights are used to generate more informed hypotheses and follow-up assays, if needed.

Finch is called to conduct full bioinformatic analyses and visualizations When provided with a disease name, Robin formulates a series of general questions about the disease pathology and queries Crow to answer each question (Supplementary section 5.2). Using the reports from Crow as context, Robin next identifies 10 potential causal disease mechanisms. For each mechanism, Robin again deploys Crow to prepare a detailed report describing an in vitro model of the disease mechanism and corresponding assay that can be used to test drug efficacy. Robin uses an LLM judge to make pairwise comparisons between reports, which are used to calculate their relative rankings (see Methods). The top-ranked in vitro model is used by Robin to define the experimental strategy for therapeutic candidate hypothesis generation.

Once an in vitro model is selected, Robin conducts a similar sequence of general literature review and hypothesis generation to propose 30 therapeutic candidates for experimental testing. Robin then queries Falcon to generate a detailed report to evaluate each candidate. These reports contain both justification for why each drug is suitable for mitigating the disease mechanism represented in the in vitro model and potential limitations the drug may pose. The drug candidates are ranked by an LLM-judged tournament according to the strength of the scientific rationale, pharmacological profile, and methodology of the supporting literature. This ranked list can then be reviewed by human scientists, and top drug candidates can be tested in the lab by executing an experimental protocol based on the assay suggested by Robin.

Once experiments are complete, the scientist uploads raw or semi-processed data and prompts Robin with a desired analysis approach, e.g., “RNA-seq differential expression analysis” or “flow cytometry". Robin then deploys Finch to carry out the desired analysis. This analysis step presents unique challenges due to the inherently ambiguous nature of biological data interpretation. For example, the gating choices in flow cytometry analysis or the filters used in RNA-seq analyses will vary between human scientists and may impact the final conclusions. Similarly, Finch’s analysis results can vary between runs, even when given identical prompts and data, due to the stochasticity of the language agent. To leverage this diversity, Robin can launch 10 Finch analysis trajectories, each of which independently analyzes the experimental data. In each trajectory, Finch executes analysis code in a Jupyter notebook and provides an interpretable and reproducible summary of its findings. After all trajectories are complete, a meta-analysis can be conducted to synthesize all outputs into a consensus-driven conclusion. In this way, Finch both explores diverse analytical trajectories while delivering highly consistent end results based on consensus (Methods) [32].

Discussion. In this report, we present Robin, a multi-agent system integrating automated hypothesis generation and experimental data analysis for scientific discovery. When tasked with identifying novel therapeutics for dAMD, Robin proposed enhancing RPE cell phagocytosis using ROCK inhibitors, and discovered ripasudil as the most potent enhancer of RPE phagocytosis among tested compounds through an iterative lab-in-the-loop discovery cycle. Ripasudil’s established safety profile and clinical approval for ocular use present a promising drug repurposing opportunity that could significantly accelerate the development pathway for dAMD treatment.

Notably, while ROCK inhibitors have been previously suggested for treatment of wet AMD and other retinal diseases of neovascularization, Robin is the first to propose their application in dry AMD for their effect on phagocytosis [35]. This approach is supported by several lines of evidence, as RPE phagocytic dysfunction is observed not only with normal aging but is pronounced in AMD patients [41, 42, 43]. In investigating the basis for Robin’s initial identification of Y-27632, we found that the literature search included a single paper demonstrating Y-27632’s ability to enhance phagocytic efficiency in “low-phagocytic” RPE cells from human donors by promoting actin polymerization [42]. Taken together, these results demonstrate Robin’s ability to effectively generate novel hypotheses by synthesizing insights already present in the scientific literature.

Beyond the specific application of dAMD, Robin addresses a broader challenge in therapeutic development. With FDA approvals stagnating at approximately 50 novel drugs annually over the past decade [44], new approaches to scale therapeutic discovery are urgently needed. As the first system to automate both literature-grounded hypothesis generation and experimental data analysis, Robin is poised to accelerate the pace of drug discovery compared to traditional approaches. We demonstrate Robin’s broader applicability by generating hypotheses for 10 additional diseases with pressing therapeutic needs (Supplementary Figures S16-S25).

By automating hypothesis generation, experimental planning, and data analysis in an integrated system, Robin represents a powerful new paradigm for AI-driven scientific discovery. This approach can be used to not only reshape therapeutic development, but fundamentally accelerate the scientific process to drive a greater understanding of the natural world.

Limitations. In its current implementation, Robin has several opportunities for continued development. For example, while Robin generates experimental outlines, it does not yet produce precise, executable protocols—future iterations aim to provide detailed methodologies that require minimal human translation for laboratory execution. The Finch data analysis agent is also heavily reliant on prompt engineering by domain experts to produce reliable analytical results. Adapting Finch to independently generate or adapt prompts to specific data modalities would enable a more autonomous discovery pipeline. Finally, while Robin uses an LLM-judged tournament to nominate therapeutic hypotheses, future work on better aligning hypothesis generation and evaluation with human scientific judgment may be helpful in more reliably producing high-quality hypotheses.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does AI-assisted research sacrifice exploration breadth for productivity gains? What human oversight must AI research systems have? Can AI systems discover fundamental improvements to their own architectures? How do clinicians calibrate trust in AI medical recommendations? What limits language model accuracy in evaluating ideas? Which reinforcement learning modifications most improve dialogue quality in language models? Why do LLM research ideation systems generate novelty but lack diversity? What prevents LLMs from applying their reasoning knowledge to improve outputs? Do language models reason through disagreement or only accommodate it? Can LLMs distinguish between linguistic form and semantic meaning? How does diversity prevent model convergence on superficial patterns?