INQUIRING LINE

Conferences ban or limit AI in peer review — do reviewers actually follow the rules, and could anyone tell?

Do reviewer rules about LLM use in peer review actually get followed?

This explores whether reviewers actually obey conference rules that ban or limit LLM use when writing peer reviews, how anyone can tell, and whether breaking those rules changes anything.


This explores whether reviewers actually obey conference rules that ban or limit LLM use when writing peer reviews, and how anyone could tell. The short answer from the corpus is: often not. The most direct evidence comes from a randomized experiment at ICML 2026. Reviewers were assigned to either a no-LLM rule or a limited-use rule, and substantial fractions broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. Before any rules existed, a population-level analysis of ICLR, NeurIPS, CoRL and EMNLP reviews estimated that between 6.5% and 16.9% of review text had been substantially modified by an LLM. The rates were higher among reviewers who reported low confidence, submitted close to the deadline, or engaged less with the paper How much peer review text shows signs of LLM modification?. That suggests rule-breaking is less about defiance and more about reviewers who are short on time or effort.

Measuring compliance is hard because LLM text is hard to spot. ICML hid instructions inside submission PDFs that would leave a telltale trace if a reviewer pasted the paper into a chatbot. This flagged 795 reviews, about 1%, and led to 497 desk rejections. The chairs themselves said the method mostly catches careless users and misses anyone who removed or rewrote the trap How many peer reviewers secretly used LLMs despite the ban?. Humans are no better detectors: ML-expert readers couldn't reliably tell LLM-written abstracts from human ones Can readers tell LLM abstracts from human ones?. So ICLR 2026 took a pragmatic route. It treated detector flags as one input for human area chairs rather than an automatic verdict, and saved hard enforcement for something checkable: fabricated references, which led to desk rejection How can conferences detect and handle LLM misuse in peer review?. In practice, enforcement is moving away from "did you use an LLM?" toward "did your work contain a verifiable error?"

Here's the twist you might not expect. In the same ICML experiment, banning versus limiting LLM use made almost no difference to paper scores, decisions, or reviewer confidence Does banning LLM use in peer review change review outcomes?. One reading is that the rules are poorly followed and also don't matter much for outcomes. The worry that LLM-assisted reviewers favor LLM-written papers also mostly dissolves on inspection. Across 125,000+ reviews, the apparent favoritism disappears once paper quality is controlled for. LLM-assisted reviewers are just more lenient toward weaker papers in general, and LLM-written papers happen to cluster among weaker submissions Do LLM reviewers actually favor LLM-written papers?. The real risks show up when the LLM does the judging. In simulations, LLM reviewers inflated scores for LLM-style prose and penalized human papers containing critical statements Do LLM reviewers favor papers written by other LLMs?. Their scores also shifted when only a paper's rhetorical framing changed and its scientific content stayed the same How much does rhetorical style shift AI review scores?.

That points to a different question than compliance: if reviewers will use LLMs anyway, can conferences channel that use well? ICLR 2025 tried giving reviewers optional, sanctioned LLM feedback on their drafts. Over a quarter of them revised, and blinded raters judged the revised reviews more specific and clearer Can LLM feedback help peer reviewers improve their own reviews?. An older study found that GPT-4 feedback overlapped with any one human reviewer about as much as two human reviewers overlapped with each other Can GPT-4 feedback match what human reviewers catch?. The corpus doesn't yet have direct data on whether sanctioned-use policies get better compliance than bans. Still, the evidence leans toward one conclusion: rules that ignore how overloaded reviewers actually work will be quietly broken, while built-in, structured LLM help may achieve what the bans were meant to protect.


Sources 10 notes

Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

How much peer review text shows signs of LLM modification?

Analysis of reviews from ICLR 2024, NeurIPS 2023, CoRL 2023, and EMNLP 2023 estimates this population share using distributional methods rather than per-review classification. Rates were higher in low-confidence, rushed, and less-engaged reviewers.

How many peer reviewers secretly used LLMs despite the ban?

Hidden-instruction watermarks planted in PDFs flagged about 1% of reviews under ICML's no-LLM rule, leading to 497 desk rejections. The chairs acknowledge the method catches mainly careless uses and misses reviewers who removed or rewrote the watermark.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Show all 10 sources
Do LLM reviewers actually favor LLM-written papers?

Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.

Do LLM reviewers favor papers written by other LLMs?

Simulated LLM reviewers gave higher scores to LLM-written papers and downrated human papers containing critical statements, while human annotators showed no such bias. The bias traces to preference for LLM writing style and aversion to critical framing.

How much does rhetorical style shift AI review scores?

Rewriting manuscripts' rhetoric while preserving scientific content moves LLM reviewer scores measurably. Evidence framing and novelty stance produce the largest contrasts; effects vary by reviewer model and interact with the paper's original quality tier.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Can GPT-4 feedback match what human reviewers catch?

A study of 3,096 Nature papers and 1,709 ICLR papers found GPT-4 matched individual reviewers' points 30.85% of the time versus 28.58% for two human reviewers. Fifty-seven percent of surveyed researchers found the feedback helpful.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.