SYNTHESIS NOTE
Topics›Domain Specialization›this note

Can LLM feedback help peer reviewers improve their own reviews?

This randomized trial tested whether optional AI-generated suggestions on review quality would prompt reviewers to revise, and whether those revisions would be more useful to authors and decision-makers.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

The paper's central claim is that optional LLM feedback on a reviewer's own review, gated by automated reliability tests, made reviews more specific and led many reviewers to revise them. The Review Feedback Agent, five LLMs with Claude Sonnet 3.5 as the backbone, ran at ICLR 2025 as a randomized control trial. The authors report that 26.6% of reviewers who received feedback updated their reviews (27% in the abstract), incorporating 12,222 suggestions. Blinded ML researchers rated the revised reviews "more informative and clearer," and reviewers who updated lengthened them by an average of 80 words. Rebuttal exchanges also ran longer in the feedback arm. The abstract's "more than 20,000" refers to randomly selected reviews. The statistics section reports feedback posted to 18,946 of 44,831 reviews (42.3%), with 3,521 selected reviews receiving none, which implies about 22,467 selected in total. That last figure is derived from the excerpt's numbers, not stated in it.

The mechanism is deliberately narrow. The agent fired once per selected review, on first submission, and posted its comments an hour later so reviewers could fix typos first. It looked for three problems: vague or generic critiques, questions that overlooked parts of the paper could answer, and unprofessional statements. Feedback went only to the reviewer and the program chairs, was not shared with authors or area chairs, and was not a factor in decisions. Reviewers were told it came from an LLM, and the system made no direct edits. Reliability tests acted as the gate: feedback was posted only if it passed every test, which kept 829 selected reviews from receiving feedback. The authors treat voluntariness as a design principle: reviewers "could opt out by ignoring the feedback." Each review took about a minute and cost about 50 cents.

Set against the nearest notes, the study asks a different question from the ICML 2026 experiment in Does banning LLM use in peer review change review outcomes?. That study measured whether reviewers obeyed rules on LLM use and found substantial noncompliance under both policies. Here no rule was involved. Uptake was the reviewer's choice, so the 27% figure measures voluntary adoption, not compliance. The agent also sits at a different point in the review chain from Can inference scaling help reviewers catch errors humans miss?. That agent audits manuscripts for flaws human reviewers missed, while this one audits reviews, so its quality measure is informativeness and specificity, not detected errors. Its gated, multi-stage design echoes Can structured pipelines make LLM novelty assessment reliable?, but it differs from Can automated review loops handle AI-generated research at scale?. aiXiv iterates review and refinement in a closed loop, whereas the ICLR agent acted only on the initial review, with no later exchange.

The excerpt does not establish that the revisions were better in any way that matters to authors. "More informative" is a blinded rating whose protocol, sample and effect sizes the excerpt does not give. The 80-word increase is conditional on updating, so it describes the revising minority, not the whole feedback arm. The discussion says feedback-arm reviewers were more likely to change their scores after rebuttal, but gives no figures, and the "significantly longer" wording comes without its test. Nothing here shows changed decisions, and the authors built the system and report its outcomes. The defensible reading is narrower than the headline: a gated, optional feedback layer at scale moved a meaningful minority of reviewers to revise, and blinded raters found the revisions clearer. Whether that improves what authors receive needs an evaluation this excerpt does not contain.

Inquiring lines that read this note 60

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI systems perform peer review as effectively as humans? What human oversight must AI research systems have? How can we detect and account for LLM involvement in academic writing? Do restrictions on reviewer LLM use actually shape peer review behavior? How can we reduce inherent biases in LLM-based evaluation judges? Do individually safe AI actions create unsafe outcomes in integrated systems? Are AI-generated articles systematically disadvantaged in search ranking and user engagement?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 64 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

feedback on peer reviews at ICLR 2025 led 27 percent of reviewers to revise and made revisions more informative — a randomized trial by its developers