How can conferences detect and handle LLM misuse in peer review?
Explores how ICLR 2026 balanced detection limitations with practical enforcement, distinguishing between acceptable LLM assistance and problematic offloading of reviewing responsibilities.
The ICLR 2026 program chairs describe an enforcement design in which LLM detection informs human judgment rather than replacing it. Their policy required that any LLM use be disclosed and that individuals remain "ultimately responsible for their contributions", with violations treated as code of ethics violations. Given the scale of over 75,000 reviews, the chairs ran two LLM content detectors on all of them and emailed area chairs about reviews that both detectors flagged as entirely LLM generated. Area chairs were asked to weigh this "as an aspect of review quality", in the same way they already weighed other signals such as excessively short reviews. On the submission side, the chairs flagged high proportions of LLM-generated content to area chairs and found one problem they judged more tractable: references to documents that did not exist, or that carried egregiously wrong bibliographic information. Papers with confirmed hallucinated references were desk rejected, with an appeal channel. The chairs say this partly explains the high desk rejection rate that year.
The reasoning is explicit about the limits. The chairs write that "systems exist for detecting LLM-generated content, their accuracy is far from perfect", and that they "did not want to penalize cases where an LLM was used to assist with writing but the judgements made in the review were reasonable and valid." The same caution shapes the reference check. The chairs call its false positive rate "significant", giving the example of a non-English title that authors had translated, so area chairs did a first human pass, the chairs checked every flagged reference themselves, and each flagged paper was reviewed by at least three humans before any rejection. The distinction the chairs draw is between LLM help with wording, which they tolerate, and LLM use "to offload their reviewing responsibilities", which they say they had to "proactively address". Hallucinated references get the firmest treatment because the chairs judged them "relatively straightforward to detect".
This is a different question from the one the ICML experiment asked. Does banning LLM use in peer review change review outcomes? tested what happens when reviewers are told to ban or limit LLM use, and found substantial noncompliance under both rules. The chairs' detector sweep works on the output after the fact, and their admission that detectors are imperfect marks the limit of that approach. The human side points the same way. Can readers tell LLM abstracts from human ones? found that readers did not reliably tell LLM-generated abstracts from human ones either, which is a human-side counterpart to the chairs' point about detection. The reference check also contrasts with Can LLM judges be fooled by fake credentials and formatting?, which lists fake references as an authority-bias lever for LLM evaluators. In this retrospective, the same kind of fabrication is instead a desk-rejection trigger.
The excerpt does not establish how many reviews or papers the flags touched, how often the detectors were right, whether flagged reviews were in fact weaker, or what happened to reviewers after the flags. The removal of "repeatedly flagged" reviewers is stated as policy, not as an outcome, and calling the approach "in line with standard practice" is the chairs' own characterization, with no outside evaluation in the excerpt. What the evidence supports is narrow. Fabricated citations can be checked with an error process that has human review and appeal built in, and text-level detection can serve as a prompt for human review. The excerpt gives no grounds for treating a detector flag as proof that a review is bad.
Inquiring lines that read this note 46
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans?- How much of ICLR 2026 peer review was already conducted by AI?
- Can LLM reviewers catch technical issues that human reviewers miss?
- How often do false positives from detection tools actually occur in peer review?
- What makes disruptive scientific work harder to publish and recognize?
- How do automated reviewers detect flaws that human experts miss in manuscripts?
- Do citation counts better capture scientific quality than publication venue tiers?
- Do preprint servers have tools to detect hidden text in submitted manuscripts?
- Why do authors submit manuscripts to venues beyond their reach?
- What role do conference organizers play in accepting problematic articles?
- How often do journal editors catch obvious textual problems before publication?
- Can LLM-generated reference reviews detect machine-written peer review submissions?
- Can readers reliably distinguish LLM-generated research writing from human writing?
- Did reviewers successfully circumvent ICML's hidden-instruction watermark detection method?
- What quality differences exist between flagged and unflagged peer reviews?
- Can text-based algorithms reliably detect LLM assistance in scientific abstracts?
- Do LLM adopters actually cite more diverse and younger research?
- Can researchers detect individual papers modified by LLMs reliably?
- What concerns does widespread LLM use raise for scientific independence?
- What methods can reliably detect LLM-generated academic papers at scale?
- Do humans and LLMs agree on novelty assessment in research?
- Can reviewer-author matching by LLM use amplify biases in acceptance decisions?
- Do peer review policies banning LLM use actually change reviewer behavior and decisions?
- What happens when conferences enforce bans or limits on reviewer LLM use?
- Can rules against undisclosed LLM use change reviewer behavior without enforcement?
- Can watermark-based detection measure true prevalence of LLM use in peer review?
- How much noncompliance occurred under ICML's limited-LLM-use policy versus the no-LLM rule?
- Why does author withdrawal authority matter for preprint accountability?
- Do reviewer rules about LLM use in peer review actually get followed?
- Do conference policies banning LLM use actually reduce AI involvement in reviews?
- Do peer reviewers actually follow policies that ban or limit their LLM use?
- How effective are journal policies restricting LLM use in peer review?
- Can peer review policies actually prevent LLM use when compliance is hard to monitor?
- Do metareviewers and regular reviewers use LLMs differently in peer review?
- How do mirror sites and shadow libraries perpetuate retracted papers?
- Why are hallucinated references easier to detect and punish than LLM-assisted writing?
- How susceptible are LLM evaluators to fake references as exploitable biases?
- How much undetected fraud exists beyond current retraction statistics?
- Can structured evaluation pipelines reduce LLM reviewer bias?
- Do LLM judges remain vulnerable to gaming when anchored to external references?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
a randomized test of the ban-or-limit rules whose enforcement the chairs' detector sweep addresses after the fact; it found substantial noncompliance under both
-
Can readers tell LLM abstracts from human ones?
Do readers with ML expertise reliably distinguish human-written, LLM-generated, and LLM-edited research abstracts? Understanding this matters for evaluating whether readers can serve as effective gatekeepers against LLM content.
the human-side counterpart to the chairs' admission that LLM-text detection is unreliable
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
contrast: fabricated references are an authority-bias lever for LLM evaluators there and a desk-rejection trigger here
-
How do detection tools shape LLM use enforcement at ICLR?
ICLR's 2026 policy uses LLM detectors to flag papers, but requires human reviewers to find concrete evidence before acting. This matters because false positives from automated tools could harm unflagged papers while creating extra work for area chairs.
qualifies: llm-detector flags act only after area chairs find concrete evidence; undisclosed heavy LLM use is also penalized, not just fabricated references
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- A Retrospective on the ICLR 2026 Review Process
- Exploring the use of AI authors and reviewers at Agents4Science
- Stop Automating Peer Review Without Rigorous Evaluation
- ICLR 2026 Response to LLM-Generated Papers and Reviews
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- On Violations of LLM Review Policies
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
Original note title
ICLR's 2026 program chairs routed llm-detector flags to area chairs as one quality input and desk-rejected papers with confirmed hallucinated references