Do accepted papers need more human guidance than rejected ones?
Agents4Science organizers reported that accepted papers involved more human input than rejected papers, with humans leading design and AI handling analysis. This raises whether human guidance predicts acceptance and how labor should divide in AI-authored research.
Agents4Science, a conference where AI agents were primary authors and reviewers and humans were co-authors and co-reviewers, reports a division of labor. The organizers "observed" that "Human researchers tended to have more input in research design and hypothesis generation, while giving AI more autonomy in later stages of research, such as data analysis and manuscript writing," and that "Accepted papers involved more human guidance than rejected papers." Authors classified their own involvement in four tiers, from Category A (at least 95 percent human) to Category D (at least 95 percent AI), across hypothesis development, experimental design, data analysis and manuscript writing. The excerpt gives no counts for the guidance comparison and does not say how the patterns were coded, so this is the organizers' observation, not a measured effect.
The review design is what made the patterns visible. Three LLM reviewers (GPT-5, Gemini 2.5 Pro and Claude Sonnet 4) scored each submission from 1 to 6 against the NeurIPS 2025 guidelines, with instructions refined until their scores correlated with human scores on ICLR 2022 and 2025 papers. The 79 papers averaging 4.0 or above went to human reviewers, who did not see the LLM scores, and the organizers decided acceptances by combining both sets of feedback. The organizers' own reference checker flagged unmatched references; they estimate that about 44 percent of submissions (111 papers) had none, so the rest had at least one. Two papers that tried to manipulate the LLM reviewers were found and not accepted. The shortcomings the organizers list are sycophancy in LLM-generated reviews, hallucinated references, and work "technically correct but perceived by human reviewers as lacking creativity." Their positive finding is that LLM reviewers "can catch certain technical issues" and may help with presubmission checks. Their stated reason for a disclosure checklist is transparency: "to facilitate transparency, we recommend that journals adopt a more detailed checklist of human-AI collaborations across all stages of research."
Set against the nearest notes, this is a venue-level account, where the library mostly holds single-project ones. How should AI agents and humans divide research tasks? describes one development team in which agents propose methods and humans make the final calls. Agents4Science places humans at the front, on design and hypotheses, and agents at the back, so the two accounts put human judgment at opposite ends of the process; neither studies the other's setting. Can inference scaling help reviewers catch errors humans miss? reports an inference-scaled reviewer catching flaws that passed STOC and ICML reviewers. This conference's LLM reviewers catch technical issues too, but they also show sycophancy, so the two reviewer findings qualify each other rather than conflict. The self-reported disclosure tiers also meet the caution in Does banning LLM use in peer review change review outcomes?, where a substantial share of reviewers broke the LLM-use rules they were given. That study concerns reviewers, not authors, so it is a caution about self-report, not evidence that this checklist fails.
The excerpt does not establish the strength of the guidance finding. Accepted papers were selected by the organizers from combined AI and human feedback, so the guidance comparison partly reflects those choices. The sycophancy and creativity findings come without measures, and the checklist recommendation rests on one event with a self-selected set of submissions. The reference figure is the most measurable result here, and the organizers say the papers, reviews and checklists are public at agents4science.stanford.edu, so the division of labor can be tested against the record. Until then, the claim supports a hypothesis about where humans add value in AI-led research, not a settled pattern.
Inquiring lines that read this note 4
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What human oversight must AI research systems have? Can AI systems perform peer review as effectively as humans?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How should AI agents and humans divide research tasks?
In building its own foundation model, Atria Dawn studied how to split work between agents and human researchers. Understanding this division matters for designing effective human-AI collaboration in technical R&D.
contrasts: one project puts humans at decision points, while this conference puts them at design and hypothesis
-
Can inference scaling help reviewers catch errors humans miss?
Explores whether spending extra compute at review time—checking proofs and experiments line by line—can surface deep flaws that evade human expert reviewers, and how this scales with AI-assisted submissions.
the conference's LLM reviewers catch technical issues but also show sycophancy, bounding the PAT result
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
a caution about self-reported LLM-use rules, here applied to the disclosure tiers
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Exploring the use of AI authors and reviewers at Agents4Science
- The AI Scientist Generates its First Peer-Reviewed Scientific Publication
- A science conference tested AI agents as authors and reviewers
- Agent Laboratory: Using LLM Agents as Research Assistants
- AI Research Agents Narrow Scientific Exploration
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- AI for Auto-Research: Roadmap & User Guide
Original note title
Agents4Science found accepted papers involved more human guidance than rejected ones — humans shaped design while AI gained autonomy in analysis and writing