Hidden traps in submitted papers caught about 1% of AI-written peer reviews, but mostly the careless ones, so the true share is unknown.
Can watermark-based detection measure true prevalence of LLM use in peer review?
This explores whether hidden watermarks planted in papers can tell us how many peer reviewers actually use LLMs, or only how many get caught using them carelessly.
This explores whether hidden watermarks planted in papers can tell us how many peer reviewers actually use LLMs, or only how many get caught. The short answer from the corpus is no. Watermarks give you a floor, not a measurement. ICML hid instructions inside submission PDFs. If a reviewer pasted the paper into an LLM, the model followed the hidden instruction and left a recognizable trace in the review. This flagged 795 reviews, about 1%, and led to 497 desk rejections. The chairs themselves say the method mostly catches careless users, and misses anyone who spotted and removed the instruction or rewrote the LLM's output How many peer reviewers secretly used LLMs despite the ban?. So the 1% counts reviewers who were careless. It does not count all the reviewers who used an LLM.
The more revealing evidence comes from a different method. ICML 2026 also ran a randomized experiment that assigned reviewers either a full ban or a limited-use policy. Substantial fractions of reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. Set that next to the watermark's 1% and the gap is the point: the watermark catches a small, visible slice of the real behavior. The same experiment found that banning versus limiting LLM use barely changed scores or decisions. That suggests the prevalence number may matter less for review outcomes than the policy debate assumes.
Other kinds of detection don't fill the gap. Readers with ML expertise couldn't reliably tell LLM-written abstracts from human ones, and tended to assume a human was involved Can readers tell LLM abstracts from human ones?. ICLR 2026 treated statistical LLM detectors as too unreliable to act on alone. It routed their flags to area chairs as one input among several, and saved hard enforcement for something checkable: fabricated references How can conferences detect and handle LLM misuse in peer review?. The pattern is to enforce only where the evidence can be confirmed, and not to treat any detector as a census.
The twist is that the hidden-prompt technique cuts both ways. The watermarks work because LLMs obey instructions buried in a document. Authors have used the same weakness: 18 arXiv manuscripts contained concealed prompts telling AI reviewers to praise the paper, a practice now classed as questionable research conduct Are hidden AI prompts in preprints a deceptive research practice?. LLM judges are also easily swayed by fake authority signals and polished formatting Can LLM judges be fooled by fake credentials and formatting?. So a measurement tool built on prompt injection depends on the same flaw that makes LLM reviewing open to manipulation. It also teaches careful reviewers what to look for, which means it is likely to undercount more over time.
If you want a truer picture of prevalence, the corpus points toward large-scale statistical and outcome-based analysis rather than catching individuals. One study of 125,000+ reviews looked at LLM-assisted reviewing across the whole review pool. It found that the apparent favoritism of LLM-assisted reviewers toward LLM-written papers disappears once you control for paper quality Do LLM reviewers actually favor LLM-written papers?. The lesson there is about method more than about cheating: even population-level signals can mislead unless you account for confounders. The corpus doesn't contain a validated true-prevalence estimate for LLM use in peer review, and that absence is itself telling.
Sources 7 notes
Hidden-instruction watermarks planted in PDFs flagged about 1% of reviews under ICML's no-LLM rule, leading to 497 desk rejections. The chairs acknowledge the method catches mainly careless uses and misses reviewers who removed or rewrote the watermark.
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
Show all 7 sources
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- On Violations of LLM Review Policies
- A Retrospective on the ICLR 2026 Review Process
- Pangram Predicts 21% of ICLR Reviews are AI-Generated