INQUIRING LINE

An AI-text detector estimated about 21% of ICLR 2026 paper reviews were fully AI-written, and over half had some AI involvement.

How much of ICLR 2026 peer review was already conducted by AI?

This explores how much of the reviewing at ICLR 2026, one of the main machine learning conferences, was actually written by AI rather than by the human reviewers assigned to it, and what that changes about the reviews.


This explores how much of the reviewing at ICLR 2026 was actually written by AI rather than by the people assigned to do it. The most direct estimate comes from Pangram Labs. It ran an AI-text detector over ICLR's public reviews and estimated that about 21% were fully AI-generated and more than half had some AI involvement How much AI content appears in peer review at ICLR?. These are a detector's predictions, not confirmed cases. Still, the picture is clear: covert AI involvement was common, not rare.

An earlier study gives a baseline, though not a perfect one. It looked at ICLR 2024, NeurIPS 2023 and other AI conferences and estimated that 6.5% to 16.9% of review text had been substantially modified by an LLM How much peer review text shows signs of LLM modification?. That study measured the whole pool of reviews statistically instead of flagging individual ones. Because the two studies counted in different ways, the numbers aren't a clean trend line. They point the same way, though. The earlier study also found who leaned on AI most: reviewers who were less confident, rushed or less engaged. So AI tends to fill in where human effort is thinnest.

The more surprising finding is what AI involvement did to the reviews. In Pangram's data, reviews with more AI text gave systematically higher scores How much AI content appears in peer review at ICLR?. AI wasn't just polishing what a human had already decided; it may have been making reviews more generous. A related problem makes this worse. AI reviewers show a 'hivemind' effect, agreeing with each other more than human reviewers do, and authors can raise AI review scores by about 0.45 points just by having a model rewrite their paper, without improving the science Can AI systems safely replace human peer reviewers?. If a fifth of reviews come from similar models, a paper is getting fewer independent opinions than its review count suggests.

ICLR 2026's organizers chose not to punish based on detector scores alone. Detector flags went to area chairs as one signal among several, and sanctions required concrete human-found evidence How do detection tools shape LLM use enforcement at ICLR?. The firm enforcement point was hallucinated references: papers with confirmed fabricated citations were rejected without full review How can conferences detect and handle LLM misuse in peer review?. Compare AAAI-26, which took the opposite approach. It openly added one labeled AI review to every fully reviewed paper, and survey respondents said they preferred those reviews on technical accuracy Did conference reviewers prefer AI reviews over human ones?.

So the honest answer is that roughly a fifth of ICLR 2026 reviews may have been written entirely by AI, and most involved it somehow. The bigger issue is that this happened without disclosure, so nobody could adjust for the inflated scores or the shared blind spots. Some argue that if AI speeds up how fast research is produced, review will have to be AI-assisted too, and the real choice is whether that happens openly and with someone accountable Can human review keep pace with AI-accelerated research generation?. Others say the fix needs authors, reviewers and venues to share responsibility, for example by letting authors rate how good a review was before they see the decision Can two-stage review and badges fix AI conference peer review?.


Sources 8 notes

How much AI content appears in peer review at ICLR?

Pangram Labs' analysis of ICLR's public review corpus estimates 21% of reviews were fully AI-generated and over half had some AI involvement. Reviews with more AI text received systematically higher scores, suggesting AI may amplify positive bias rather than just rephrase human judgment.

How much peer review text shows signs of LLM modification?

Analysis of reviews from ICLR 2024, NeurIPS 2023, CoRL 2023, and EMNLP 2023 estimates this population share using distributional methods rather than per-review classification. Rates were higher in low-confidence, rushed, and less-engaged reviewers.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

How do detection tools shape LLM use enforcement at ICLR?

ICLR 2026 uses LLM detection tools only to triage papers for area chairs, who must find concrete evidence before sanctions are applied. This human gate protects against false positives by shifting costs from paper rejections to reviewer time.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Show all 8 sources
Did conference reviewers prefer AI reviews over human ones?

At AAAI-26, every full-review paper received one labeled AI review generated by a multi-stage pipeline. Survey respondents reported preferring these AI reviews over human reviews on dimensions like technical accuracy, though the preference size and statistical strength were not disclosed.

Can human review keep pace with AI-accelerated research generation?

The PAT framework argues that accepting AI-driven research output commits us structurally to AI-assisted verification, not as option but necessity. A taxonomy of four collaboration levels—from author tools to reviewer augmentation—provides the governance scaffolding to manage this transition while keeping humans accountable.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.