How many peer reviewers secretly used LLMs despite the ban?
ICML used hidden watermarks in submission PDFs to detect LLM-written reviews submitted under a no-LLM policy. The question explores whether 795 flagged reviews represent the true scale of LLM use or only careless violations.
ICML's 2026 program chairs report that they caught LLM-written reviews by planting instructions in submission PDFs, and that the enforcement ended in 497 desk rejections. Of the reviews written under Policy A, the rule that no LLMs may be used, 795 (about 1% of all reviews) from 506 reviewers were flagged, and a human checked each flagged instance. The rejections follow from the rule's consequence: a flagged review by a reciprocal reviewer got that reviewer's own submission rejected, so the 497 papers correspond to 398 reciprocal reviewers. The chairs also removed 51 reviewers who had used LLMs in more than half of their reviews. The post was corrected on one point: the remaining 108 flagged reviewers were not reciprocal reviewers of active submissions.
The mechanism, which the chairs trace to recent work by Rao, Kumar, Lakkaraju, and Shah, turns the reviewer's own tool into the detector. The chairs built a dictionary of 170,000 phrases and sampled two per paper, a pair they say is less likely than one in ten billion. Each PDF carried instructions "visible only to an LLM" to include those two phrases in the review. A human reader would not see them, but a model reading the PDF would, and the chairs say the watermark "would subtly influence any review produced via an LLM." A review containing both phrases is therefore a fingerprint of a model that read the paper. Generic AI-text detectors were not used. The chairs report a family-wise false positive rate of 0.0001 and say every flag was inspected to rule out reviews that merely mentioned the watermark.
This is different evidence from the randomized experiment in Does banning LLM use in peer review change review outcomes?, which compared scores and decisions across policies and found noncompliance under both. The ICML post counts only what the watermark caught under the no-LLM rule, so it measures detected violations, not how often reviewers used LLMs. The same channel appears in Can LLM judges be fooled by fake credentials and formatting?, where hidden text steers a judge; here the hidden instruction is used as a detector instead. The excerpt also notes that the method "is not a difficult measure to circumvent," particularly once publicly known, which it was for almost the entire review period. Success rates were "over 80% for most models" in pre-deadline tests, but "not always."
The excerpt does not establish how common LLM use was among reviewers without the rule, or among those on Policy B, where limited use was allowed, so it cannot say whether 795 reviews is a large or small share of LLM-assisted reviewing. It also does not assess quality: "we are not making a judgment call about the quality of flagged reviews or the reviewers' intentions." The watermark misses anyone who removed it, rewrote the review, or used a model that ignored it. The implication, at the strength the evidence allows, is that 795 is a floor for what one method caught under one policy, not a prevalence estimate. Measuring noncompliance cleanly would take a design like the experiment's, not the watermark alone.
Inquiring lines that read this note 11
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do restrictions on reviewer LLM use actually shape peer review behavior?- Do peer review policies banning LLM use actually change reviewer behavior and decisions?
- What happens when conferences enforce bans or limits on reviewer LLM use?
- Can rules against undisclosed LLM use change reviewer behavior without enforcement?
- Can watermark-based detection measure true prevalence of LLM use in peer review?
- How much noncompliance occurred under ICML's limited-LLM-use policy versus the no-LLM rule?
- Do reviewer rules about LLM use in peer review actually get followed?
- Do peer reviewers actually follow policies that ban or limit their LLM use?
- How effective are journal policies restricting LLM use in peer review?
- Can peer review policies actually prevent LLM use when compliance is hard to monitor?
Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
Same policy design, different evidence: the experiment measured noncompliance under both policies, while this post counts only watermark catches under the no-LLM rule.
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
Same channel, opposite purpose: hidden instructions steer LLM judges in that note and here serve as a detector of LLM use in a review.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- On Violations of LLM Review Policies
- A Retrospective on the ICLR 2026 Review Process
- ICLR 2026 Response to LLM-Generated Papers and Reviews
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Stop Automating Peer Review Without Rigorous Evaluation
- Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
Original note title
ICML says hidden-instruction watermarks flagged 795 reviews under its no-LLM rule and led to 497 desk rejections — possibly only the most careless uses