Over half of reviewers say they already use AI tools, so how would an author tell which review was machine-written?
How often do researchers suspect peer reviews are written by AI?
This explores how common it is for researchers to believe the peer reviews they receive were written by AI, and what the corpus can say about the AI use behind that suspicion.
This explores how often researchers suspect that a peer review was written by AI. The collection has no survey that measures suspicion directly. It does show what that suspicion would be responding to, and the gap is large. In a Frontiers survey of 1,645 researchers, 53% of reviewers said they already use AI tools, and among early-career researchers the figure was 87%. Most use it to draft reports or summarize papers How widely do peer reviewers actually use AI tools?. So the useful question is less whether authors suspect AI and more how much of what they read was shaped by it. On the survey's numbers, that's likely more than half of reviews.
A second finding explains why suspicion is hard to settle. When writers get AI-drafted text, they edit it only 23% of the time, and their edits leave the text about 96% the same as before Do writers actually edit AI-generated text before publishing?. That study wasn't about reviewers. If reviewers behave the same way, though, a review that was 'assisted' may read almost exactly like one that was generated, and the line an author would want to draw gets blurry. One sign authors might pick up on: AI reviewers agree with each other far more than human reviewers do, a 'hivemind' effect. They are also easy to game. Rewriting a paper's text, with no change to the science, raised AI scores by 0.45 points Can AI systems safely replace human peer reviewers?.
The clearest evidence that people assume AI is reading comes from authors, not reviewers. Eighteen arXiv manuscripts were found with hidden instructions telling AI reviewers to give positive assessments Are hidden AI prompts in preprints a deceptive research practice?. Nobody plants a hidden prompt unless they expect a model to read the paper. That's suspicion turned into strategy. A survey of 230 publications describes this as an arms race in which AI-written papers, AI review, manipulation, and defenses each push the others forward Does AI create a coupled arms race in research production and review?.
Some venues have stopped guessing and made AI involvement visible. At ICLR 2025, a randomized trial gave reviewers optional feedback from an AI agent, and 27% of them revised their reviews to be more specific Can LLM feedback help peer reviewers improve their own reviews?. One proposal argues that problems with AI conference reviewing are shared by authors, reviewers, and venues. It suggests letting authors rate review quality before they see the decision Can two-stage review and badges fix AI conference peer review?. Others argue that if AI speeds up how fast research is produced, AI-assisted review stops being optional and becomes necessary Can human review keep pace with AI-accelerated research generation?.
The surprise is that the most useful AI reviewer may be the opposite of the generic one authors worry about. PAT is an agent that spends extra computing time checking proofs and experiments line by line. It found serious flaws in papers that human reviewers had passed at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. So 'was this written by AI?' may matter less than 'did anyone, human or AI, actually check the work?' For a direct measure of how often authors suspect AI, the collection doesn't have one yet.
Sources 9 notes
Frontiers' May-June 2025 survey of 1,645 researchers found 53% of reviewers use AI tools, with adoption reaching 87% among early-career researchers. Most use AI for drafting reports or summarizing findings, and researchers express desire for clearer policies to guide more advanced applications.
Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
Show all 9 sources
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
The PAT framework argues that accepting AI-driven research output commits us structurally to AI-assisted verification, not as option but necessity. A taxonomy of four collaboration levels—from author tools to reviewer augmentation—provides the governance scaffolding to manage this transition while keeping humans accountable.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- Towards Automating Scientific Review with Google's Paper Assistant Tool