INQUIRING LINE

Reviewers who use AI most often ask it to draft reports and summarize papers, though task-level data is thin.

What specific tasks do reviewers use AI for most often?

This explores what peer reviewers actually use AI for day to day (drafting, summarizing, checking, and so on), and what the collection can and can't tell you about that.


This explores what peer reviewers actually use AI for day to day: which parts of reviewing they hand off, as opposed to whether they use it at all. The collection gives only a partial answer. The clearest data point is Frontiers' 2025 survey of 1,645 researchers. It found that 53% of reviewers use AI tools, rising to 87% among early-career researchers. The most common uses were **drafting the review report** and **summarizing a paper's findings** How widely do peer reviewers actually use AI tools?. Those are writing and reading tasks, not judgment tasks. The same researchers said they wanted clearer policies before moving on to anything more ambitious. Beyond that one survey, the collection has no detailed breakdown of tasks by frequency.

Other studies show where drafting help shades into something larger. Pangram Labs analyzed ICLR's public reviews and estimated that 21% were fully AI-generated and more than half had some AI involvement How much AI content appears in peer review at ICLR?. The surprising part is that reviews with more AI text gave systematically higher scores. That suggests 'help me write this up' may be changing the verdict, not just the wording. It fits a separate finding that AI reviewers tend to agree with each other more than humans do, and that rewording a paper can raise AI scores without improving the science Can AI systems safely replace human peer reviewers?.

The rules around AI use also seem to matter less than you might expect. At ICML 2026, a randomized experiment either banned LLM use or allowed limited use. Scores and decisions barely changed, and many reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. So reviewers' actual habits may not match what venues permit or what surveys record.

The tasks that look most promising are not the ones reviewers report doing most. One is **checking the reviewer's own review**. In an ICLR 2025 trial, AI feedback on draft reviews led 27% of reviewers to revise, and blinded raters judged the revised reviews more specific and clearer Can LLM feedback help peer reviewers improve their own reviews?. Another is **line-by-line error-hunting**. An agentic system that spends extra compute checking proofs and experiments found flaws in STOC and ICML papers that human reviewers had missed Can inference scaling help reviewers catch errors humans miss?. Some venues now supply AI reviews officially. AAAI-26 attached a labeled AI review to every full-review paper, and survey respondents rated those reviews higher on technical accuracy Did conference reviewers prefer AI reviews over human ones?.

The takeaway: reviewers mostly use AI for the writing side of the job, summarizing and drafting. The evidence suggests AI is more useful on the checking side, auditing math and stress-testing a reviewer's own critique. Few reviewers report using it that way so far.


Sources 7 notes

How widely do peer reviewers actually use AI tools?

Frontiers' May-June 2025 survey of 1,645 researchers found 53% of reviewers use AI tools, with adoption reaching 87% among early-career researchers. Most use AI for drafting reports or summarizing findings, and researchers express desire for clearer policies to guide more advanced applications.

How much AI content appears in peer review at ICLR?

Pangram Labs' analysis of ICLR's public review corpus estimates 21% of reviews were fully AI-generated and over half had some AI involvement. Reviews with more AI text received systematically higher scores, suggesting AI may amplify positive bias rather than just rephrase human judgment.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Show all 7 sources
Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Did conference reviewers prefer AI reviews over human ones?

At AAAI-26, every full-review paper received one labeled AI review generated by a multi-stage pipeline. Survey respondents reported preferring these AI reviews over human reviews on dimensions like technical accuracy, though the preference size and statistical strength were not disclosed.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.