AI Research Agents Narrow Scientific Exploration

Paper · arXiv 2605.27905 · Published May 27, 2026
Correct but Not Understood

Abstract AI research agents now support large-scale AI-assisted scientific discovery. We examine whether AI-generated ideas broaden scientific exploration or primarily reinforce existing work. Using five agent frameworks and five large language models, we generate 219,655 ideas for different scientific fields. Across experiments, four consistent patterns emerge. First, AI-generated ideas are more concentrated than human-authored papers within the same research area. Second, they remain much closer to starting literature than later human follow-on work does. Third, AIgenerated ideas align less with future human research. Last, AI-generated ideas are located in lower-impact regions of the historical scientific landscape. Overall, current AI research agents appear better suited to local elaboration than to broadening scientific exploration.

Keywords: Artificial intelligence, Agentic AI, Scientific discovery, Large language models, Science of science

Introduction. Recent advances in AI research agents have raised the possibility of automating scientific discovery. These agents can now conduct literature reviews, generate research ideas, plan experiments, run code, write papers, and iteratively explore and refine scientific hypotheses [1–4]. Importantly, these AI research agent frameworks are explicitly designed to encourage exploratory scientific ideation. Their prompts and reasoning procedures often instruct agents to generate novel, high-impact, and unconventional ideas rather than simple extensions of prior work [1–3]. As such systems become increasingly capable and accessible, they may fundamentally reshape the process of scientific discovery. Yet the ability to generate research ideas at scale does not necessarily imply broader scientific exploration. Scientific discovery often depends on moving beyond established directions, searching less familiar regions, and recombining prior knowledge in non-routine ways [5, 6]. Existing evaluations of AI research agents mainly assess whether individual ideas are interesting, novel, feasible, or executable [7, 8], but reveal much less about how repeated AI-assisted ideation may shape the broader landscape of scientific exploration. This raises a broader question: do AI research agents broaden scientific exploration? To study this question, we construct research areas from the scientific literature across broad fields and use AI research agents to generate ideas from shared seed papers within each research area. Specifically, using papers published between 2020 and 2025 from the Semantic Scholar Academic Graph as seed literature, we use advanced AI research-agent frameworks, including AIScientist [1], ResearchAgent [2], AgentLaboratory [3], and Co-Scientist [9], together with five LLMs to generate complete scientific research ideas, including both research questions and methods. In total, we analyze 219,655 valid AI-generated research ideas spanning 12 broad scientific fields and 155 research areas. Throughout the ideation process, all evaluated AI agent frameworks are explicitly instructed to explore novel research directions beyond the seed literature. The AI

Method. 2 Generating Scientific Ideas with AI Agents Define Research Areas.

We begin by constructing research areas from the scientific literature across major fields of science. We collect papers from the Semantic Scholar Academic Graph1, together with their reference and citation information. The resulting corpus spans 12 fields, including Medicine, Biology, Engineering, Chemistry, Computer Science, Environmental Science, Materials Science, Physics, Mathematics, Economics, Business, and Sociology. Each paper record includes its abstract, publication year, primary field, and citation links. We use papers published before 2020 to construct research areas. Within each scientific field, we identify research areas using bibliographic coupling [10]. Specifically, papers are represented by their citation profiles and clustered according to bibliographiccoupling similarity. The final identified research areas span topics including cryo-electron microscopy, pancreatic cancer treatment, offshore wind power, heavy-ion physics, and microbiome community assembly. Detailed construction procedures are provided in the Supplementary Methods (see SM S1.1).

Scientific Idea generation.

We next use AI research agents to generate new scientific ideas from prior literature. For each identified research area, We repeatedly sample seed-paper sets from papers published between 2020 and 2025 to initialize AI ideation. Each seed set contains five papers: one anchor paper together with four related papers from the same research area, selected using citation. We use five seed papers because most evaluated AI research agent frameworks are constrained by context-window limitations of current LLMs. These seed papers define a coherent research topic and provide a common starting literature context for idea generation. During idea generation, the evaluated agent frameworks can further search a locally deployed Semantic Scholar database to retrieve additional relevant papers, allowing them to expand the literature context beyond the initial seed set. To preserve the historical setting, agents are only allowed to retrieve papers that were available when the seed papers were published. We evaluate five representative AI research-agent frameworks: a Zero-shot baseline, AIScientist [1], ResearchAgent [2], AgentLaboratory [3], and Co-Scientist [9]. These frameworks represent several major designs for AI research agents, including iterative self-reflection, multi-stage planning and validation, multi-agent deliberation, and tournament-based hypothesis evolution. All evaluated agent frameworks, except the Zero-shot baseline, can retrieve additional papers from the locally deployed Semantic Scholar database. Table 1 summarizes the evaluated frameworks and their corresponding agentic mechanisms. Importantly, across all evaluated frameworks, the ideation prompts explicitly encourage exploration beyond the seed literature. The Zero-shot baseline asks the model to propose a novel research idea. AIScientist emphasizes generating “novel” and “high-impact” ideas through iterative selfreflection and revision. ResearchAgent encourages innovative method design during multi-stage planning and validation. Agent Laboratory explicitly instructs agents to expand upon the literature and generate ideas that are “very innovative and unlike anything seen before.” Co-Scientist generates and refines hypotheses through comparison, and explicitly rewards novelty at each round. Supplementary Methods provides the full prompts and detailed agentic design of each evaluated framework (see SM S1.2). Each AI agent framework needs to be paired with an LLM. We evaluate five LLMs: four openweight models ranging from 8B to 35B parameters, Gemma-4-31B-IT [11], Llama-3.1-8B [12], Hermes-4-14B [13], and Qwen3-35B-A3B [14], together with OpenAI’s GPT-5.4 [15]. Across the five agent frameworks and five LLMs, we bootstrap seed-paper sets from the identified research areas. In total, our analysis uses 219,655 valid AI-generated ideas generated from 155 research areas spanning 12 broad scientific fields, obtained from 232,800 generation runs.

Discussion. 4.1 Exploration Breadth: AI Ideas Are More Concentrated Than Human Papers We examine the exploration breadth of AI-generated ideas within each identified research area. As a comparison, we also measure the exploration breadth of the human-authored seed papers. To ensure a fair comparison, we randomly sample one seed paper for each AI-generated idea so that the human and AI collections contain the same number of ideas. Figure 2a–b shows a consistent pattern across the five agent frameworks and five LLMs. AI-generated ideas within the same research area are more similar to one another than humanauthored papers from those same areas. Averaged across agent frameworks, the breadth is 0.554 for AI-generated ideas and 0.599 for human-authored papers, representing a 7.5% lower exploration breadth for AI-generated ideas. Panels c–d further show that different LLMs and agent frameworks often explore overlapping regions of the same research area. Within the same research area, the average exploration breadth between ideas generated by different agent frameworks is 0.572, while that between ideas generated by different LLMs is 0.570. These averages remain lower than the human same-area baseline. Fieldlevel analyses show the same pattern across 11 of the 12 broad scientific fields, with Mathematics being the only exception (see SM S2.1 and table S5). We observe the same pattern using an alternative centroid-based measure of exploration breadth. For each research area, AI-generated ideas lie closer to their area centroids than humanauthored papers do, again indicating a more concentrated exploration pattern (see SM S2.2 and table S6).

4.2 Exploration Distance: AI Ideas Stay Close to Their Starting Literature We next examine the exploration distance of AI-generated ideas to assess whether they move beyond the seed literature used to initialize ideation or instead remain locally anchored to it. As a comparison, we examine follow-on human-authored papers that directly cite at least one of the seed papers [22]. These follow-on papers represent subsequent human research emerging from the same research topic. Figure 3 compares the distributions of exploration distance for AI-generated ideas and followon human papers across four consecutive years (2020→2021 through 2023→2024). Across all four years, AI-generated ideas remain closer to the seed literature than subsequent human research. In every year, the AI distributions are shifted toward smaller exploration distances, whereas followon human papers exhibit broader distributions extending to substantially larger distances. On average, across years, exploration distance is 0.322 for AI-generated ideas, compared with 0.410 for follow-on human papers. Field-level analyses reveal the same pattern across all twelve broad scientific fields, with the difference remaining statistically significant in every field (see SM S2.1 and table S5). This result suggests that, although all evaluated AI agent frameworks can search for and retrieve relevant literature from the Semantic Scholar database, AI-generated ideas remain largely confined to local exploration.

4.3 Frontier Alignment: AI Ideas Are Less Aligned with Future Research Frontiers We next examine the frontier alignment of AI-generated ideas to assess whether they align with future directions of scientific research. As a comparison, we measure the frontier alignment for follow-on human-authored papers that directly cite at least one of the seed papers. To ensure a fair comparison, these follow-on papers are excluded from the next-year corpus when constructing the frontier keyword set. AI-generated ideas and follow-on human papers contain a comparable number of extracted keywords on average (11.80 versus 11.88).

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do LLM research ideation systems generate novelty but lack diversity? Does AI-assisted research sacrifice exploration breadth for productivity gains? What human oversight must AI research systems have? Can AI systems discover fundamental improvements to their own architectures? How do clinicians calibrate trust in AI medical recommendations?