How much AI content appears in peer review at ICLR?
A detection company scanned 70,000 ICLR reviews to estimate how prevalent AI-generated text is in peer review. Understanding this prevalence matters for assessing review quality and integrity in academic publishing.
Pangram Labs, a company that sells AI-text detection, analyzed all 19,000 papers and 70,000 reviews from ICLR using the public OpenReview record. Its headline figure is that "21%, or 15,899 reviews, were fully AI-generated," and that "over half of the reviews had some form of AI involvement." On the submission side it reports that 61% of papers were "mostly human-written," that several hundred were fully AI-generated, and that 9% had over 50% AI content. These are the detector's estimates, produced by the company that built the detector. The excerpt gives its false positive rate as 1 in 10,000 on test documents and 1 in 100,000 on held-out arXiv papers, and compares that to a drug test.
The score pattern is the sharpest claim. Pangram finds that "the more AI is present in a review, the higher the score is," and reads this as reviewers outsourcing the judgment itself rather than rephrasing their own view: if AI were only a framing device, average scores on AI and human reviews should match. Sycophancy is offered as one explanation for the positive bias. The company also argues that length no longer signals care. AI reviews run long and carry "filler content" with "low information density," a property it attributes to Shaib et al.'s Measuring AI Slop in Text. It reverses an earlier finding that LLM judges favor their own outputs: here, "the more AI-generated text present in a submission, the worse the reviews are."
Against the nearest notes, this is an observational prevalence scan, not a test of policy. Does banning LLM use in peer review change review outcomes? randomizes the rules and measures what changed. Pangram's data cannot separate AI use from the reviewers who choose it, so it sizes the behavior rather than its effect. The sycophancy account is one candidate mechanism for the score link. How much does rhetorical style shift AI review scores? points to another: scores move with framing even when content is held fixed, so AI-written reviews could score higher on presentation alone. The excerpt tests neither. Its detection concern also rests on a machine check, since Can people reliably spot content made by AI? suggests readers cannot reliably spot this by eye. That fits Pangram's call for conferences to adopt AI detection.
What the excerpt does not establish: it gives no method for how the detector labels a review "fully AI-generated" beyond the error rates, and it does not say how partial AI involvement was scored, so the "over half" figure rests on a loose category. The paper figures also undercount by the company's own caveat, since some fully AI-generated papers were removed before the scan. The score association is a correlation across reviews, and the excerpt does not say whether it controls for paper quality, topic or venue. The 21% is therefore best read as a detector-based estimate with a published error rate, and the score link as an association worth testing rather than a settled cause. The policy conclusion, that fully AI-generated reviews should be sanctioned, is the company's and the excerpt presents it as its own position.
Inquiring lines that read this note 6
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans? What are the real-world consequences of AI citation hallucinations? How do educators verify student capability when AI can produce indistinguishable work?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
a randomized ICML test of peer-review AI rules; this scan measures prevalence, not policy effects.
-
Does polished writing actually signal better quality work?
When evaluators judge applications and manuscripts, does rhetorical sophistication predict merit, or does it distract from verifiable evidence of competence and rigor?
shares the warning that surface fluency misleads judgments of merit; Pangram adds that AI-heavy reviews score higher.
-
How much does rhetorical style shift AI review scores?
When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.
a competing mechanism for the score link: framing alone can shift scores.
-
Can people reliably spot content made by AI?
This systematic review of 30 studies asks whether human judgment can distinguish AI-generated text, images, and voice from human-created content, and whether detection accuracy has improved as AI becomes more realistic.
explains why a machine detector, not reader judgment, carries the prevalence claim.
-
How much peer review text shows signs of LLM modification?
Researchers analyzed AI conference reviews to estimate what fraction might have been substantially altered by large language models. Understanding this helps clarify how AI tools are entering academic peer review.
qualifies: a different corpus-level estimate (6.5–16.9%) for other venues and years, using 'substantially LLM-modified' rather than fully AI-generated
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Stop Automating Peer Review Without Rigorous Evaluation
- AI Now Writes as Many Online Articles as Humans
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Exploring the use of AI authors and reviewers at Agents4Science
- The AI Scientist Generates its First Peer-Reviewed Scientific Publication
Original note title
Pangram predicts 21% of ICLR reviews are fully AI-generated — and the more AI in a review, the higher its score