When an AI writes a peer review, what mistakes do the people reading it actually catch, and what slips past them?
What specific errors did participants report finding in the AI-generated reviews?
This explores what kinds of mistakes people found when they read peer reviews written by AI. The retrieved notes don't include a participant-level list of those errors, so this answer covers what the collection does document about where AI reviewing goes wrong.
This explores what kinds of mistakes people found when they read peer reviews written by AI. The direct answer is that none of the notes retrieved here breaks down the specific errors that participants reported. Treat what follows as the surrounding picture rather than a catalogue of complaints. That picture is still worth having, because the errors that are documented are not the ones you might expect.
The most concrete error in the retrieved notes runs the other way: it is an AI-written paper that got through human review. When Sakana AI's fully automated system produced a paper that met the acceptance threshold at an ICLR 2025 workshop, the authors went back afterwards and found a citation error that the human reviewers had missed. They also judged that none of their three submissions was ready for a main conference Can AI-generated papers pass peer review undetected?. That fits a broader finding: people generally can't tell AI-generated content from human work at better than chance Can people reliably spot content made by AI?. Fluent, confident output tends to hide its own mistakes, and the cases where it matters most are the ones that average accuracy scores miss Why do confident wrong answers hide in standard accuracy metrics?.
When the AI does the reviewing, the documented weaknesses are mostly systemic. A single review might read well; the problem shows up across many reviews. AI reviewers agree with each other far more than human reviewers do, which the authors call a 'hivemind' effect. Simply rewording a paper's text, without changing the science, raised AI review scores by about half a point Can AI systems safely replace human peer reviewers?. This is related to the way models over-trust answers they generated themselves Why do models trust their own generated answers?. An AI reviewer can be persuaded by polished writing in the same way.
The strongest counterpoint is that AI reviewing improves when it is given time to check the work. PAT, an agentic reviewer that works through proofs and experiments line by line, found flawed proofs and broken experiments in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. At ICLR 2025, AI feedback on human reviews led 27% of reviewers to revise them, and independent raters judged the revised reviews more specific and clearer Can LLM feedback help peer reviewers improve their own reviews?. Across these results, AI does worse as a one-shot judge and better as a careful checker or an assistant to a human reviewer.
If you're looking for the actual list of complaints from participants in a specific pilot, such as a conference survey of authors who received labeled AI reviews, the retrieved notes don't contain it. That detail would have to come from the original survey report.
Sources 7 notes
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.
Medical triage, legal interpretation, and financial planning show a consistent pattern: surface heuristics conflict with unstated constraints, producing fluent confident errors that concentrate in rare cases where harm occurs. Aggregate accuracy masks these failures because overall performance looks strong.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
LLMs exhibit structural bias toward validating their own outputs because high-probability generated answers feel more correct during evaluation. Comparing answers against broader alternatives breaks this self-agreement loop.
Show all 7 sources
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content