When peer reviewers were randomly told to ban AI or allow limited use, scores barely moved, and many broke the rule anyway.
How do disclosure rules in peer review actually change reviewer behavior and scores?
This explores whether rules about AI use in peer review (banning it, allowing some, or requiring reviewers to say when they used it) actually change how reviewers behave and how papers get scored, and what nearby research on AI disclosure suggests about why.
This explores whether rules about AI use in peer review change what reviewers do and how papers score. The corpus has one direct test, and it's a surprising one. At ICML 2026, reviewers were randomly assigned either a ban on LLM use or permission for limited use. The difference in paper scores, accept/reject decisions, and reviewer confidence came out close to zero Does banning LLM use in peer review change review outcomes?. Large shares of reviewers also broke whichever rule they were given. So in the one place it has been measured, the rule changed behavior less than the people writing it probably hoped, and it barely changed outcomes. Beyond that experiment, the collection has little that studies peer review disclosure rules directly. The rest of this answer draws on nearby work.
If reviewers use AI anyway, rules matter less than what the AI does once it's in the loop. LLM reviewers move their scores when a paper's rhetoric changes, even if the science stays the same. How confidently evidence is framed and how boldly novelty is claimed cause the biggest swings How much does rhetorical style shift AI review scores?. A reviewer who quietly breaks a ban may be passing that sensitivity to style into their score. The rule doesn't remove that effect. It just hides it. Some researchers are taking the opposite approach: build AI review into the process openly, with structure, rather than trying to police it. Examples include breaking novelty checks into separate steps Can structured pipelines make LLM novelty assessment reliable?, agents that spend extra computing effort checking proofs line by line Can inference scaling help reviewers catch errors humans miss?, and closed review-and-revise loops for AI-written research Can automated review loops handle AI-generated research at scale?.
Research on disclosure outside peer review helps explain why disclosure is a weak lever. When an article says AI helped write it, both human and LLM raters score it lower, but only slightly: less than 0.15 points on a 7-point scale Does disclosing AI assistance make readers trust articles less?. The penalty grows in personal, relationship-based writing, where readers see AI use as a breach of social expectations How does revealing AI authorship change reader trust?. Peer review is closer to the technical end of that range. Readers and writers also disagree about when disclosure is needed. Readers want it more, especially when AI text goes into the final work unchanged Do readers and writers differ on AI disclosure necessity?. That gap is one reason reviewers might under-report their own AI use.
The less obvious finding is that labels can shift human and AI judges in opposite directions. When told a human wrote a piece that broke a writing rule, AI evaluators went easier on it, while human judges got stricter Do authorship labels change how AI judges evaluate rule violations?. Disclosure effects can also wear off. People's initial bias against an AI partner reversed once they saw its results repeatedly, but only when they got that feedback Does revealing AI identity help or hurt user trust?. Together these suggest a disclosure rule's effect depends on who reads the label and whether anyone later sees how the reviews turned out. Simply having the rule on the books doesn't determine much. It also suggests a better question than "ban or allow?": what feedback would let a venue tell whether AI-assisted reviews are actually better or worse?
Sources 10 notes
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Rewriting manuscripts' rhetoric while preserving scientific content moves LLM reviewer scores measurably. Evidence framing and novelty stance produce the largest contrasts; effects vary by reviewer model and interact with the paper's original quality tier.
A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.
Show all 10 sources
Both human raters (n=1,970) and LLM raters (n=2,520) scored an identical news article lower when it included an AI disclosure statement, but the penalty was small—less than 0.15 points on a 7-point scale.
A study of 261 readers found that disclosing AI authorship consistently lowered perceived trustworthiness, caring, and likability, with the steepest drops in interpersonal writing like personal interaction. Readers saw AI as incapable of genuine empathy, viewing its use as a violation of social expectations.
A 727-person vignette study found readers consistently rated AI disclosure as more necessary than writers did. Disclosure seemed most necessary when AI text was directly incorporated and irreplaceable, while writer effort had no effect on these judgments.
AI models chose a rule-breaking lipogram 35 percentage points more often when told a human wrote it, while human judges chose it 20 points less in that condition. The shift suggests AI may relax standards for human work while humans anchor to objective compliance.
Users initially avoid AI partners when identity is revealed, but this preference reverses after repeated interactions with visible results. The learning mechanism—observing consistent outcomes—is essential; disclosure without feedback produces no calibration.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- What Influences Readers' and Writers' Perceived Necessity of AI Disclosure?
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- LLM-REVal: Can We Trust LLM Reviewers Yet?