INQUIRING LINE

When an AI writes a peer review, what mistakes do the people reading it actually catch, and what slips past them?

What specific errors did participants report finding in the AI-generated reviews?

This explores what kinds of mistakes people found when they read peer reviews written by AI. The retrieved notes don't include a participant-level list of those errors, so this answer covers what the collection does document about where AI reviewing goes wrong.


This explores what kinds of mistakes people found when they read peer reviews written by AI. The direct answer is that none of the notes retrieved here breaks down the specific errors that participants reported. Treat what follows as the surrounding picture rather than a catalogue of complaints. That picture is still worth having, because the errors that are documented are not the ones you might expect.

The most concrete error in the retrieved notes runs the other way: it is an AI-written paper that got through human review. When Sakana AI's fully automated system produced a paper that met the acceptance threshold at an ICLR 2025 workshop, the authors went back afterwards and found a citation error that the human reviewers had missed. They also judged that none of their three submissions was ready for a main conference Can AI-generated papers pass peer review undetected?. That fits a broader finding: people generally can't tell AI-generated content from human work at better than chance Can people reliably spot content made by AI?. Fluent, confident output tends to hide its own mistakes, and the cases where it matters most are the ones that average accuracy scores miss Why do confident wrong answers hide in standard accuracy metrics?.

When the AI does the reviewing, the documented weaknesses are mostly systemic. A single review might read well; the problem shows up across many reviews. AI reviewers agree with each other far more than human reviewers do, which the authors call a 'hivemind' effect. Simply rewording a paper's text, without changing the science, raised AI review scores by about half a point Can AI systems safely replace human peer reviewers?. This is related to the way models over-trust answers they generated themselves Why do models trust their own generated answers?. An AI reviewer can be persuaded by polished writing in the same way.

The strongest counterpoint is that AI reviewing improves when it is given time to check the work. PAT, an agentic reviewer that works through proofs and experiments line by line, found flawed proofs and broken experiments in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. At ICLR 2025, AI feedback on human reviews led 27% of reviewers to revise them, and independent raters judged the revised reviews more specific and clearer Can LLM feedback help peer reviewers improve their own reviews?. Across these results, AI does worse as a one-shot judge and better as a careful checker or an assistant to a human reviewer.

If you're looking for the actual list of complaints from participants in a specific pilot, such as a conference survey of authors who received labeled AI reviews, the retrieved notes don't contain it. That detail would have to come from the original survey report.


Sources 7 notes

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can people reliably spot content made by AI?

A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.

Why do confident wrong answers hide in standard accuracy metrics?

Medical triage, legal interpretation, and financial planning show a consistent pattern: surface heuristics conflict with unstated constraints, producing fluent confident errors that concentrate in rare cases where harm occurs. Aggregate accuracy masks these failures because overall performance looks strong.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Why do models trust their own generated answers?

LLMs exhibit structural bias toward validating their own outputs because high-probability generated answers feel more correct during evaluation. Comparing answers against broader alternatives breaks this self-agreement loop.

Show all 7 sources
Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.