SYNTHESIS NOTE
Topics›Domain Specialization›this note

How can conferences detect and handle LLM misuse in peer review?

Explores how ICLR 2026 balanced detection limitations with practical enforcement, distinguishing between acceptable LLM assistance and problematic offloading of reviewing responsibilities.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

The ICLR 2026 program chairs describe an enforcement design in which LLM detection informs human judgment rather than replacing it. Their policy required that any LLM use be disclosed and that individuals remain "ultimately responsible for their contributions", with violations treated as code of ethics violations. Given the scale of over 75,000 reviews, the chairs ran two LLM content detectors on all of them and emailed area chairs about reviews that both detectors flagged as entirely LLM generated. Area chairs were asked to weigh this "as an aspect of review quality", in the same way they already weighed other signals such as excessively short reviews. On the submission side, the chairs flagged high proportions of LLM-generated content to area chairs and found one problem they judged more tractable: references to documents that did not exist, or that carried egregiously wrong bibliographic information. Papers with confirmed hallucinated references were desk rejected, with an appeal channel. The chairs say this partly explains the high desk rejection rate that year.

The reasoning is explicit about the limits. The chairs write that "systems exist for detecting LLM-generated content, their accuracy is far from perfect", and that they "did not want to penalize cases where an LLM was used to assist with writing but the judgements made in the review were reasonable and valid." The same caution shapes the reference check. The chairs call its false positive rate "significant", giving the example of a non-English title that authors had translated, so area chairs did a first human pass, the chairs checked every flagged reference themselves, and each flagged paper was reviewed by at least three humans before any rejection. The distinction the chairs draw is between LLM help with wording, which they tolerate, and LLM use "to offload their reviewing responsibilities", which they say they had to "proactively address". Hallucinated references get the firmest treatment because the chairs judged them "relatively straightforward to detect".

This is a different question from the one the ICML experiment asked. Does banning LLM use in peer review change review outcomes? tested what happens when reviewers are told to ban or limit LLM use, and found substantial noncompliance under both rules. The chairs' detector sweep works on the output after the fact, and their admission that detectors are imperfect marks the limit of that approach. The human side points the same way. Can readers tell LLM abstracts from human ones? found that readers did not reliably tell LLM-generated abstracts from human ones either, which is a human-side counterpart to the chairs' point about detection. The reference check also contrasts with Can LLM judges be fooled by fake credentials and formatting?, which lists fake references as an authority-bias lever for LLM evaluators. In this retrospective, the same kind of fabrication is instead a desk-rejection trigger.

The excerpt does not establish how many reviews or papers the flags touched, how often the detectors were right, whether flagged reviews were in fact weaker, or what happened to reviewers after the flags. The removal of "repeatedly flagged" reviewers is stated as policy, not as an outcome, and calling the approach "in line with standard practice" is the chairs' own characterization, with no outside evaluation in the excerpt. What the evidence supports is narrow. Fabricated citations can be checked with an error process that has human review and appeal built in, and text-level detection can serve as a prompt for human review. The excerpt gives no grounds for treating a detector flag as proof that a review is bad.

Inquiring lines that read this note 46

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI systems perform peer review as effectively as humans? How can we detect and account for LLM involvement in academic writing? Do restrictions on reviewer LLM use actually shape peer review behavior? How do hallucinated citations emerge in AI scholarly output? How reliably can humans and AI detectors identify machine-generated text? What limits language model accuracy in evaluating ideas? What external process records should verify agent behavior and benchmark claims? How can we reduce inherent biases in LLM-based evaluation judges? What are the real-world consequences of AI citation hallucinations? What prevents LLMs from applying their reasoning knowledge to improve outputs? What governance mechanisms can effectively constrain widely deployed AI systems?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 78 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

ICLR's 2026 program chairs routed llm-detector flags to area chairs as one quality input and desk-rejected papers with confirmed hallucinated references