Can AI-generated papers pass peer review undetected?
Explores whether end-to-end AI-generated manuscripts can clear human double-blind review at academic workshops, and what acceptance rates reveal about reviewer capability to distinguish AI from human work.
Sakana AI reports that one of the three papers its AI Scientist-v2 generated passed double-blind peer review at an ICLR 2025 workshop. The papers were "entirely generated end-to-end by AI, without any modifications from humans": the system proposed the hypothesis, designed and ran the experiments, analyzed the data, and wrote every word of the manuscript. The human team chose only the broad topic and which three papers to submit. The accepted paper, "Compositional Regularization: Unexpected Obstacles in Enhancing Neural Network Generalization", reports a negative result and averaged 6.33 from reviewers. Sakana says this ranks "approximately 45% of all submissions" and sits above the workshop's average acceptance threshold. Only one of the three was accepted.
The evidence comes from a designed experiment rather than an ordinary submission. Reviewers were told that 3 of 43 papers might be AI-generated, but not which ones. Under a protocol agreed in advance, any accepted AI paper would be withdrawn and desk-rejected, and kept off OpenReview's public forum, because the AI and scientific communities have not decided whether AI-generated manuscripts belong in the same venues. The organizers also skipped a meta-review because they already knew of the experiment, so the 6.33 is a set of reviewer scores with no final decision behind it. The stated reason for withdrawal is about norms, not about the quality of the reviews.
Read against the library, this is the verification gap in concrete form. The note on Can AI verify research outputs as fast as it generates them? argues that generation runs ahead of proof. Here the artifacts cleared a human review, and the builders' own reading then found a citation error (an LSTM network credited to Goodfellow (2016) rather than Hochreiter and Schmidhuber (1997)) and judged that none of the three met the bar for an ICLR main-track paper. That fits the worry in Does polished writing actually signal better quality work?, though the excerpt shows only scores and does not show that reviewers were swayed by polish. It also contrasts with Can inference scaling help reviewers catch errors humans miss?: that note describes a machine check finding flaws that passed human review, while this excerpt does not say whether the reviewers saw the errors the authors later found. The ICML experiment in Does banning LLM use in peer review change review outcomes? measured reviewer behavior under LLM rules. This excerpt measures how reviewers score papers when AI authorship is possible but unidentified.
The excerpt does not establish much beyond its own account. It reports one workshop, three submissions and one acceptance, with no reviewer comments, no independent replication, and no analysis of how reviewers reached their scores. The acceptance-rate comparison it offers (20-30% at main conferences, 60-70% at workshops) is Sakana's own framing, and the builders call the work preliminary. The implication, at the strength the evidence allows, is narrow: an end-to-end AI pipeline can produce a manuscript that clears one workshop's double-blind review at a score its builders describe as above threshold. It does not show that the work would clear a main-track bar, which Sakana's own review says it did not, and it does not show that review scores track merit. The forecast that such systems will reach top journals is Sakana's prediction, not a finding.
Inquiring lines that read this note 74
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans?- What specific errors did participants report finding in the AI-generated reviews?
- Can multi-stage AI review pipelines catch scientific flaws better than simple language models?
- Did adding AI reviews actually change peer review decisions or paper outcomes?
- Should rhetorical polish in AI reviews be separated from actual technical accuracy?
- Can human reviewers detect when papers have been rewritten by AI?
- Do AI reviews depend more on writing style than scientific merit?
- Should AI research papers require dedicated automated review systems instead of existing journals?
- Can workshop acceptance rates reliably measure AI research quality compared to main conferences?
- Can automated review systems catch deep methodological flaws or only surface issues?
- Could AI improve peer review rigor and catch human-missed errors?
- Could AI feedback work as a substitute for human peer review entirely?
- Can polished AI text fool both reviewers and detection methods?
- Can computational inference scaling catch flaws that human expert reviewers miss?
- How often do false positives from detection tools actually occur in peer review?
- What makes disruptive scientific work harder to publish and recognize?
- How can arXiv and journals scale quality control for AI-generated research?
- Can institutional statements alone correct misconceptions from unreviewed papers?
- Do AI-generated research reviews score papers higher than human reviewers do?
- Does AI content in reviews correlate with differences in paper quality control?
- Can human reviewers reliably detect AI-written peer review text by sight?
- Can machine review catch flaws in AI-generated work that humans miss?
- How often do AI systems produce papers with undetected factual errors?
- How do automated reviewers detect flaws that human experts miss in manuscripts?
- Should AI-generated papers use specialized review venues instead of traditional journals?
- Why do peer reviewers favor novel ideas that later fail in execution?
- Are refereed venues also overwhelmed by AI-generated low-quality submissions?
- Can AI reviewers detect deep theoretical flaws that human experts miss?
- Does rhetorical presentation bias reviewers against substantive scientific contributions?
- Does rhetorical quality in reviews influence paper acceptance scores more than content?
- Can agentic AI systems catch flaws in manuscripts that human reviewers consistently miss?
- Could hidden prompts be inserted during review and removed before publication?
- Why do authors submit manuscripts to venues beyond their reach?
- Could automated review systems handle AI-generated research at scale?
- Which feedback loops in AI-mediated review remain unmeasured or rarely observed directly?
- Can humans reliably detect whether research text was written by AI?
- Do shortened peer review timelines correlate with lower quality publications?
- How often do journal editors catch obvious textual problems before publication?
- Does an automated reviewer's output actually match human review accuracy?
- How do AI-generated papers perform when submitted to real conferences?
- What limitations did the authors acknowledge about their automated reviewer?
- Can feeding review scores back into idea generation improve research quality?
- Why do researchers resist using AI for peer review specifically?
- Can AI systems write and review research while operating outside traditional PDF constraints?
- How should hiring and promotion weigh AI-inflated research output?
- Can automated reviewers actually handle the review load AI creates?
- How does opaque AI methodology undermine peer review and reproducibility?
- Can technical accuracy in AI training data replace human review before publication?
- Can traditional complexity measures still signal research quality in AI-era papers?
- How do citation errors in AI-generated papers differ from human hallucinations?
- Do surface phrases reliably identify unedited machine-generated scholarship?
- Can novelty filters using literature search prevent AI-generated research from duplicating prior work?
- Can AI systems distinguish fabricated papers from legitimate research?
- How often do AI book summaries fabricate details when spot-checks are random?
- Does performing the source verification work create meaningful engagement with ideas?
- How often do fabricated sources in AI output escape citation checking?
- Will automated paper generation enable large-scale P-hacking and data dredging?
- What role should humans play in reviewing and approving AI-generated research?
- What error rates appear in AI research output when humans do not verify results?
- Are paper mills using NHANES data to automate single-factor research?
- What would a practical reviewer checklist for autonomous research systems need to include?
- Can disclosure alone ensure independent verification of AI-assisted mathematical work?
- Can mathematics remain trustworthy when results bypass peer review entirely?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can AI verify research outputs as fast as it generates them?
Research suggests AI systems produce plausible findings rapidly but struggle to verify them at the same pace. This creates a bottleneck in verification across all research stages. Understanding this gap matters for assessing when AI assistance is reliable versus risky.
the builders' own main-track review is the verification step that rejected all three after the workshop accepted one
-
Does polished writing actually signal better quality work?
When evaluators judge applications and manuscripts, does rhetorical sophistication predict merit, or does it distract from verifiable evidence of competence and rigor?
the same concern about how reviewers judge AI text, here seen only through scores
-
Can inference scaling help reviewers catch errors humans miss?
Explores whether spending extra compute at review time—checking proofs and experiments line by line—can surface deep flaws that evade human expert reviewers, and how this scales with AI-assisted submissions.
contrast: a paper passing human review still carried errors the authors caught on reading, and no machine check is described
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
both are controlled tests of peer review under AI authorship, differing in whether reviewers knew which papers were AI-generated
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The AI Scientist Generates its First Peer-Reviewed Scientific Publication
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Exploring the use of AI authors and reviewers at Agents4Science
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
Original note title
Sakana AI reports one of three fully AI-generated papers passed double-blind ICLR 2025 workshop review, then withdrew it