Are peer reviews flagged as AI-assisted actually worse than unflagged ones, or does the flag say less than you'd expect?
What quality differences exist between flagged and unflagged peer reviews?
This explores whether peer reviews flagged as LLM-written or LLM-assisted are actually worse than unflagged ones, and what the collection can say about that gap.
This explores whether peer reviews flagged as LLM-written or LLM-assisted are actually worse than unflagged ones. The collection has no study that directly compares the quality of flagged and unflagged reviews. It does contain several studies that come at the question from the side, and together they suggest the flag itself says less than you might expect.
Start with how flags are used in practice. At ICLR 2026, program chairs treated LLM-detector flags as one piece of evidence for area chairs to weigh, not as an automatic verdict. The detectors were too imperfect to trust alone, and several human review steps were added to catch false positives. The one thing they enforced strictly was fabricated references, because those can be checked How can conferences detect and handle LLM misuse in peer review?. A flag, then, is a reason to look closer, not proof that a review is bad. A randomized experiment at ICML 2026 points the same way: banning LLM use versus allowing limited use made almost no difference to paper scores, decisions or reviewer confidence. Many reviewers also broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. If AI use were badly hurting reviews, a ban should have changed the results, and it didn't.
Where a difference does appear, it is subtler than "AI reviews are sloppy." A study of more than 125,000 reviews found that LLM-assisted reviewers seemed to favor LLM-written papers. That apparent bias disappeared once paper quality was taken into account. The real pattern was that LLM-assisted reviewers are more lenient toward weaker papers in general, and LLM-written papers tend to be among the weaker submissions Do LLM reviewers actually favor LLM-written papers?. That points to a specific possible weakness: AI-assisted reviews may be too soft on weak papers. Another study adds a risk at the level of the whole system. AI reviewers agree with each other more than human reviewers do, which the authors call a "hivemind" effect, and their scores can be raised by simply rewording a paper without improving its science Can AI systems safely replace human peer reviewers?. A single AI-assisted review may look fine, while many of them together reduce the range of independent judgments that peer review depends on.
The surprise is that AI involvement can also make reviews better. In a randomized trial at ICLR 2025, reviewers could choose to receive feedback on their draft reviews from Claude-based agents. Twenty-seven percent of them revised their reviews, and blinded raters judged the revised versions more informative and clearer Can LLM feedback help peer reviewers improve their own reviews?. A detector might well flag those improved reviews. Purpose-built tools go further still: a structured pipeline for judging novelty matched human reviewers' reasoning 86% of the time Can structured pipelines make LLM novelty assessment reliable?, and an agentic reviewer that spends extra computation on checking caught errors in proofs and experiments that had passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?.
The takeaway is that "flagged versus unflagged" may be the wrong way to split reviews. The more useful questions are how AI was used (to polish and check a review, or to replace the reviewer's own judgment) and why it was used. One model suggests that overloaded reviewers produce worse reviews whether or not AI is involved, and that this overload feeds a cycle of rising submissions Does peer review quality collapse under submission overload?. A position paper argues for letting authors rate how useful a review was and rewarding thorough reviewers Can two-stage review and badges fix AI conference peer review?. That would measure review quality directly, instead of relying on a detector as a stand-in.
Sources 9 notes
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Show all 9 sources
A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
A two-journal model shows that rising submissions overtax unpaid reviewers, forcing journals to recruit less qualified reviewers or overload existing ones, which drops review accuracy and incentivizes authors to submit more speculatively, driving submissions higher. The mechanism is structural but its empirical strength remains to be measured.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?