SYNTHESIS NOTE
Topics›Domain Specialization›this note

Do accepted papers need more human guidance than rejected ones?

Agents4Science organizers reported that accepted papers involved more human input than rejected papers, with humans leading design and AI handling analysis. This raises whether human guidance predicts acceptance and how labor should divide in AI-authored research.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

Agents4Science, a conference where AI agents were primary authors and reviewers and humans were co-authors and co-reviewers, reports a division of labor. The organizers "observed" that "Human researchers tended to have more input in research design and hypothesis generation, while giving AI more autonomy in later stages of research, such as data analysis and manuscript writing," and that "Accepted papers involved more human guidance than rejected papers." Authors classified their own involvement in four tiers, from Category A (at least 95 percent human) to Category D (at least 95 percent AI), across hypothesis development, experimental design, data analysis and manuscript writing. The excerpt gives no counts for the guidance comparison and does not say how the patterns were coded, so this is the organizers' observation, not a measured effect.

The review design is what made the patterns visible. Three LLM reviewers (GPT-5, Gemini 2.5 Pro and Claude Sonnet 4) scored each submission from 1 to 6 against the NeurIPS 2025 guidelines, with instructions refined until their scores correlated with human scores on ICLR 2022 and 2025 papers. The 79 papers averaging 4.0 or above went to human reviewers, who did not see the LLM scores, and the organizers decided acceptances by combining both sets of feedback. The organizers' own reference checker flagged unmatched references; they estimate that about 44 percent of submissions (111 papers) had none, so the rest had at least one. Two papers that tried to manipulate the LLM reviewers were found and not accepted. The shortcomings the organizers list are sycophancy in LLM-generated reviews, hallucinated references, and work "technically correct but perceived by human reviewers as lacking creativity." Their positive finding is that LLM reviewers "can catch certain technical issues" and may help with presubmission checks. Their stated reason for a disclosure checklist is transparency: "to facilitate transparency, we recommend that journals adopt a more detailed checklist of human-AI collaborations across all stages of research."

Set against the nearest notes, this is a venue-level account, where the library mostly holds single-project ones. How should AI agents and humans divide research tasks? describes one development team in which agents propose methods and humans make the final calls. Agents4Science places humans at the front, on design and hypotheses, and agents at the back, so the two accounts put human judgment at opposite ends of the process; neither studies the other's setting. Can inference scaling help reviewers catch errors humans miss? reports an inference-scaled reviewer catching flaws that passed STOC and ICML reviewers. This conference's LLM reviewers catch technical issues too, but they also show sycophancy, so the two reviewer findings qualify each other rather than conflict. The self-reported disclosure tiers also meet the caution in Does banning LLM use in peer review change review outcomes?, where a substantial share of reviewers broke the LLM-use rules they were given. That study concerns reviewers, not authors, so it is a caution about self-report, not evidence that this checklist fails.

The excerpt does not establish the strength of the guidance finding. Accepted papers were selected by the organizers from combined AI and human feedback, so the guidance comparison partly reflects those choices. The sycophancy and creativity findings come without measures, and the checklist recommendation rests on one event with a self-selected set of submissions. The reference figure is the most measurable result here, and the organizers say the papers, reviews and checklists are public at agents4science.stanford.edu, so the division of labor can be tested against the record. Until then, the claim supports a hypothesis about where humans add value in AI-led research, not a settled pattern.

Inquiring lines that read this note 4

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What human oversight must AI research systems have? Can AI systems perform peer review as effectively as humans?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 69 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Agents4Science found accepted papers involved more human guidance than rejected ones — humans shaped design while AI gained autonomy in analysis and writing