An AI-text detector estimated about 21% of ICLR 2026 paper reviews were fully AI-written, and over half had some AI involvement.
How much of ICLR 2026 peer review was already conducted by AI?
This explores how much of the reviewing at ICLR 2026, one of the main machine learning conferences, was actually written by AI rather than by the human reviewers assigned to it, and what that changes about the reviews.
This explores how much of the reviewing at ICLR 2026 was actually written by AI rather than by the people assigned to do it. The most direct estimate comes from Pangram Labs. It ran an AI-text detector over ICLR's public reviews and estimated that about 21% were fully AI-generated and more than half had some AI involvement How much AI content appears in peer review at ICLR?. These are a detector's predictions, not confirmed cases. Still, the picture is clear: covert AI involvement was common, not rare.
An earlier study gives a baseline, though not a perfect one. It looked at ICLR 2024, NeurIPS 2023 and other AI conferences and estimated that 6.5% to 16.9% of review text had been substantially modified by an LLM How much peer review text shows signs of LLM modification?. That study measured the whole pool of reviews statistically instead of flagging individual ones. Because the two studies counted in different ways, the numbers aren't a clean trend line. They point the same way, though. The earlier study also found who leaned on AI most: reviewers who were less confident, rushed or less engaged. So AI tends to fill in where human effort is thinnest.
The more surprising finding is what AI involvement did to the reviews. In Pangram's data, reviews with more AI text gave systematically higher scores How much AI content appears in peer review at ICLR?. AI wasn't just polishing what a human had already decided; it may have been making reviews more generous. A related problem makes this worse. AI reviewers show a 'hivemind' effect, agreeing with each other more than human reviewers do, and authors can raise AI review scores by about 0.45 points just by having a model rewrite their paper, without improving the science Can AI systems safely replace human peer reviewers?. If a fifth of reviews come from similar models, a paper is getting fewer independent opinions than its review count suggests.
ICLR 2026's organizers chose not to punish based on detector scores alone. Detector flags went to area chairs as one signal among several, and sanctions required concrete human-found evidence How do detection tools shape LLM use enforcement at ICLR?. The firm enforcement point was hallucinated references: papers with confirmed fabricated citations were rejected without full review How can conferences detect and handle LLM misuse in peer review?. Compare AAAI-26, which took the opposite approach. It openly added one labeled AI review to every fully reviewed paper, and survey respondents said they preferred those reviews on technical accuracy Did conference reviewers prefer AI reviews over human ones?.
So the honest answer is that roughly a fifth of ICLR 2026 reviews may have been written entirely by AI, and most involved it somehow. The bigger issue is that this happened without disclosure, so nobody could adjust for the inflated scores or the shared blind spots. Some argue that if AI speeds up how fast research is produced, review will have to be AI-assisted too, and the real choice is whether that happens openly and with someone accountable Can human review keep pace with AI-accelerated research generation?. Others say the fix needs authors, reviewers and venues to share responsibility, for example by letting authors rate how good a review was before they see the decision Can two-stage review and badges fix AI conference peer review?.
Sources 8 notes
Pangram Labs' analysis of ICLR's public review corpus estimates 21% of reviews were fully AI-generated and over half had some AI involvement. Reviews with more AI text received systematically higher scores, suggesting AI may amplify positive bias rather than just rephrase human judgment.
Analysis of reviews from ICLR 2024, NeurIPS 2023, CoRL 2023, and EMNLP 2023 estimates this population share using distributional methods rather than per-review classification. Rates were higher in low-confidence, rushed, and less-engaged reviewers.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
ICLR 2026 uses LLM detection tools only to triage papers for area chairs, who must find concrete evidence before sanctions are applied. This human gate protects against false positives by shifting costs from paper rejections to reviewer time.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
Show all 8 sources
At AAAI-26, every full-review paper received one labeled AI review generated by a multi-stage pipeline. Survey respondents reported preferring these AI reviews over human reviews on dimensions like technical accuracy, though the preference size and statistical strength were not disclosed.
The PAT framework argues that accepting AI-driven research output commits us structurally to AI-assisted verification, not as option but necessity. A taxonomy of four collaboration levels—from author tools to reviewer augmentation—provides the governance scaffolding to manage this transition while keeping humans accountable.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026