A randomized test found banning AI from peer review barely moved scores, and many reviewers broke the ban anyway, so do the rules work?
Can peer review policies actually prevent LLM use when compliance is hard to monitor?
This explores whether conference rules banning or limiting LLM use in peer review actually change reviewer behavior and review outcomes, given that organizers can't reliably see who is using AI.
This explores whether conference rules about LLM use in peer review do anything when organizers can't reliably tell who is following them. The most direct evidence suggests a ban changes less than you'd expect, and the reason isn't only that people cheat. At ICML 2026, reviewers were randomly assigned either a ban on LLMs or permission for limited use. Paper scores, accept/reject decisions and reviewer confidence came out almost the same under both rules, and a large share of reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. So the policy didn't reliably control behavior, and the behavior it was meant to control didn't clearly change outcomes either.
Detection doesn't close the gap. ICML hid instructions in submission PDFs as a kind of watermark: a reviewer who pasted the paper into an LLM might get output that gave them away. This flagged 795 reviews, about 1% of the total, and led to 497 desk rejections. The chairs acknowledge it mostly caught careless users and missed anyone who removed the watermark or rewrote the output How many peer reviewers secretly used LLMs despite the ban?. Human judgment won't fill the hole. Even ML-literate readers can't reliably tell LLM-written research abstracts from human ones Can readers tell LLM abstracts from human ones?. One feared harm also turns out to be partly an illusion. Across more than 125,000 reviews, LLM-assisted reviewers seemed to favor LLM-written papers, but the effect disappears once paper quality is held constant Do LLM reviewers actually favor LLM-written papers?.
What seems to work is enforcing something you can actually check, rather than trying to catch LLM use itself. ICLR 2026 treated detector flags as one input for area chairs, not an automatic verdict. It did desk-reject papers with confirmed fabricated references, because a reference that doesn't exist can be proven How can conferences detect and handle LLM misuse in peer review?. The same logic shows up in work on keeping LLM judges honest: run the checks nobody can argue with first, and use planted cases as alarms. Neither depends on anyone confessing how they worked Can deterministic checks protect LLM judges from failure?. Fake references matter for another reason too. LLM evaluators reliably score text higher when it carries fake citations or polished formatting Can LLM judges be fooled by fake credentials and formatting?.
The surprising turn is that the more productive policy may be to channel LLM use instead of banning it. At ICLR 2025, a randomized trial offered reviewers optional, Claude-based feedback on their drafts. Over a quarter revised their reviews, and blinded raters judged the revisions more specific and informative Can LLM feedback help peer reviewers improve their own reviews?. Structured pipelines can also make narrow tasks more reliable. One that pulls out a paper's claims, retrieves related work and then compares the two reached 86% reasoning alignment with human reviewers on judging novelty Can structured pipelines make LLM novelty assessment reliable?. Put together, the evidence moves the question away from whether you can stop reviewers using LLMs. A better question is which parts of a review can be verified, and where officially sanctioned AI help improves reviews more than hidden AI help harms them.
Sources 9 notes
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Hidden-instruction watermarks planted in PDFs flagged about 1% of reviews under ICML's no-LLM rule, leading to 497 desk rejections. The chairs acknowledge the method catches mainly careless uses and misses reviewers who removed or rewrote the watermark.
Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
Show all 9 sources
Research identifies four mechanical safeguards: ordering unarguable checks before contestable ones, measuring correctness against human labels, hiding test data from proposers, and using planted cases as alarms. None requires the LLM itself to verify compliance.
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Stop Automating Peer Review Without Rigorous Evaluation
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- On Violations of LLM Review Policies
- A Retrospective on the ICLR 2026 Review Process