Exploring the use of AI authors and reviewers at Agents4Science

Paper · arXiv 2511.15534 · Published November 19, 2025
Domain Specialization in LLMs

There is growing interest in using AI agents for scientific research, yet fundamental questions remain about their capabilities as scientists and reviewers. To explore these questions, we organized Agents4Science, the first conference in which AI agents serve as both primary authors and reviewers, with humans as co-authors and co-reviewers. Here, we discuss the key learnings from the conference and their implications for human-AI collaboration in science.

Introduction. Artificial intelligence (AI) are no longer just tools for science – they now act as ‘co-scientists’, participating in all stages of research design and analysis. Traditionally, researchers begin with a welldefined question, such as predicting protein structures from amino acid sequences, and then develop or apply AI tools (like AlphaFold) to solve that specific problem. Over the past year, researchers have increasingly begun to use AI as a co-scientist to participate in a broader range of scientific activities including hypothesis generation, experimental design, and paper writing [1-5]. These AI co-scientists are powered by advances in AI agents—autonomous systems built on top of large language models (LLMs) that can use existing tools, access external databases, and search through scientific literature.

While there are promising examples of AI co-scientists designing nanobodies and generating experimentally validated hypotheses, this remains an emerging frontier [1,2]. Many fundamental questions are still open: How creative are AI scientist agents? How should human researchers collaborate with them? How capable are LLMs at reviewing scientific work? These questions are difficult to study because journals and conferences currently prohibit AI co-authors and LLM reviewers, and researchers often hide how they use AI [5].

To address this gap, we organized Agents4Science, the first conference where AI agents served as both authors and reviewers, with humans as co-authors and additional reviewers. This event provided an opportunity to explore the future of AI-driven science.

Method. Agents4Science solicited AI-led research papers across all domains of science. Each paper was primarily authored by AI agents, meaning that AI played the role of first author, as in a conventional paper, and should have made substantial contributions to project planning, execution, and writing. Humans could be co-authors. Each submission was required to complete two mandatory checklists. The first checklist was adapted from the NeurIPS conference standards and addressed general methodological and ethical considerations related to the research. The second was an Agents4Science-specific checklist designed to ensure transparency by requiring authors to disclose the extent and nature of AI involvement throughout the research process.

We required authors to disclose the extent of AI involvement in their research using a four-tier system: Category A (≥95% human contribution), Category B (50–95% human), Category C (50–95% AI), and Category D (≥95% AI). Authors reported these classifications across four key stages of the scientific process: hypothesis development, experimental design and implementation, data analysis and interpretation of results, and manuscript writing.

We created three LLM reviewers using GPT-5, Gemini2.5 Pro, and Claude Sonnet 4. Each submission was separately assessed by these three LLM reviewers, which gave each paper a rating on a scale of 1(negative) to 6(positive), using the NeurIPS 2025 reviewing guidelines as a rubric. Papers with an average score of 4.0 or above advanced to the next stage (79 papers). Human experts were invited to review these 79 papers without access to the LLM reviews. Finally, the organizers determined accepted papers, spotlight presentations, and best paper awards by synthesizing feedback from both AI and human reviewers.

To calibrate the LLM reviewers, we used anonymized papers from the ICLR 2022 and ICLR 2025 editions, along with their acceptance decisions and review scores. We then iteratively refined the instructions to the LLM reviewers to optimize correlations between average human scores and the LLM reviewer scores. All three LLM reviewers were given the same prompt and instructions to score each submission. The human reviewers were given the same reviewing guidelines.

In addition to the LLM reviewers, we also implemented several additional checks to assess the quality of each submission. First, given potential concerns about LLM-hallucinated references [9], we created and deployed an automated reference checking system. For each reference in every submission, our system extracted the reference title and other available information and conducted a web search to look for matches. If no match was found, this reference was flagged as a potential hallucination, and a notification of concern was posted on OpenReview, tagging a sample of the references in question. Figure 1d quantifies the reference hallucination rate across all the submissions. We estimate that approximately 44% of submissions have no hallucinated references (111 papers), and the other papers have one or more references flagged as problematic. This suggests that reference hallucination is still a widespread issue and requires careful checking by human co-authors.

In addition to reference checking, we implemented an automated pipeline to detect system abuse. This system scanned the entire submission to look for potential prompt injections and other adversarial instructions attempting to manipulate the LLM reviewers. Using the checker, we detected two papers that attempted to manipulate the LLM reviewers; they were not accepted.

Discussion. As AI agents become more deeply integrated into scientific research, it is essential for the community to take an evidence-based and transparent approach to understanding both their strengths and limitations as co-researchers and co-reviewers. The Agents4Science Conference represents a timely step in this direction. By making all submitted papers, reviews, checklists, and conference recordings publicly available at https://agents4science.stanford.edu/, the conference provides a rich dataset for investigating how AI agents contribute to science, where they fall short, and how humans collaborate with them.

The accepted papers illustrate the promising potential of AI agents as co-scientists and co-reviewers across a wide range of domains—from engineering and medicine to the social sciences. In particular, LLM reviewers can catch certain technical issues in manuscripts and may serve as useful tools for presubmission checks or as assistants to human reviewers [10]. At the same time, the conference surfaced important shortcomings. These include instances of sycophancy in LLM-generated reviews, hallucinated references, and research outputs that were technically correct but perceived by human reviewers as lacking creativity. Understanding such limitations was a key motivation for organizing the conference.

We also observed notable patterns in human–AI collaboration. Human researchers tended to have more input in research design and hypothesis generation, while giving AI more autonomy in later stages of research, such as data analysis and manuscript writing. Accepted papers involved more human guidance than rejected papers. Developing best practices and norms for collaboration between humans and AI is a critical area for future work. To facilitate transparency, we recommend that journals adopt a more detailed checklist of human-AI collaborations across all stages of research, similar to the one used by Agents4Science.

Conclusion. Overall, Agents4Science is a jumping-off point for the scientific community to transparently explore effective practices, ethical norms, and collaborative models in this rapidly evolving era of AI-augmented research.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems perform peer review as effectively as humans? How can we detect and account for LLM involvement in academic writing? Do restrictions on reviewer LLM use actually shape peer review behavior? How do hallucinated citations emerge in AI scholarly output? How reliably can humans and AI detectors identify machine-generated text? What limits language model accuracy in evaluating ideas?