A randomized test at a major AI conference found reviewer AI rules barely moved scores, and many reviewers broke them anyway.
What happens when reviewers use AI tools against journal policy?
This explores what actually happens to review quality, review outcomes, and the wider publishing system when peer reviewers quietly use AI tools that the venue has banned or restricted. Most of the evidence here comes from AI conferences rather than journals, but the dynamics carry over.
This explores what happens when reviewers use AI against the rules: does it change decisions, who gets hurt, and how does the rest of the system react? The most direct evidence is surprising. In a randomized experiment at ICML 2026, some reviewers were told not to use LLMs and others were allowed limited use. The rule made almost no difference to scores, accept/reject decisions, or reviewer confidence, and large shares of reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. So rule-breaking is common, and at least in measurable outcomes it doesn't obviously tilt results. Policy may be shaping what reviewers admit to more than what they actually do.
The bigger effect may be on authors, who now assume their paper will be read by a machine whatever the policy says. Researchers found hidden instructions in 18 arXiv manuscripts telling AI reviewers to give positive assessments. The authors were betting that reviewers would paste the paper into a chatbot Are hidden AI prompts in preprints a deceptive research practice?. That bet makes sense, because AI evaluators are easy to sway. They give higher scores to text with fake references or polished formatting Can LLM judges be tricked without accessing their internals?, and simply rewriting a paper's wording, with no change to the science, raised AI review scores by about half a point Can AI systems safely replace human peer reviewers?. Hidden AI use by reviewers therefore creates an opening that a dishonest author can exploit. A banned tool is still a target.
There is also a quieter problem: sameness. AI reviewers tend to agree with each other more than human reviewers do, a 'hivemind' effect Can AI systems safely replace human peer reviewers?. Peer review relies on several independent readers catching different things. If two of three 'independent' reviewers are secretly running the same model, the paper effectively gets one opinion counted twice, and nobody knows it.
Enforcement is hard in both directions. Accusing someone of using AI is unreliable. One study found that text accused of being AI-written lacked the features that actually separate AI writing from human writing, so accusations can end up wrongly punishing honest writers Do unfounded AI accusations harm human writers instead?. Openly admitting AI use also costs something: disclosure labels lower ratings slightly, though consistently Does disclosing AI assistance make readers trust articles less?. That study looked at news articles, not reviews, but it suggests why people might hide AI use instead of declaring it. A survey of 230 publications describes the whole situation as a linked arms race. More AI-produced papers lead to more automated review, which leads to manipulation, then defenses, then evasion, with each move setting off the next Does AI create a coupled arms race in research production and review?.
Some of the material suggests a ban may be the wrong tool altogether. A carefully designed AI reviewer that checks proofs line by line found serious errors that had already passed human review at top venues Can inference scaling help reviewers catch errors humans miss?. Proposals such as review-and-revise loops with built-in defenses against hidden prompts Can automated review loops handle AI-generated research at scale?, or systems that let authors rate how good their reviews were and reward thorough reviewers Can two-stage review and badges fix AI conference peer review?, treat the problem as one of design and incentives rather than policing. What the collection doesn't yet have is direct evidence from journals specifically, or data on whether hidden AI use has changed the fate of individual papers.
Sources 10 notes
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Accused comments lack features that distinguish AI text from human writing, suggesting accusations function as gatekeeping rather than detection. This inverts the AI-as-perpetrator framing, placing harm at the receiving side through reader skepticism.
Show all 10 sources
Both human raters (n=1,970) and LLM raters (n=2,520) scored an identical news article lower when it included an AI disclosure statement, but the penalty was small—less than 0.15 points on a 7-point scale.
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026