Can two-stage review and badges fix AI conference peer review?
A position paper diagnoses AI conference review failures across authors, reviewers, and venues, proposing staged author feedback on reviews and a reviewer reward system. Does this approach actually reduce bias and improve review quality?
The paper's authors argue that the peer review crisis at major AI conferences cannot be placed on reviewers alone. They name three parties, authors, reviewers and "the System" (the venue and OpenReview), and say all three "share responsibility for the current problems." Author misconduct, they add, "can only be addressed through policy enforcement and detection tools," so they focus on reviewer accountability: a two-stage bi-directional review in which authors rate reviews, and a systematic reviewer reward system.
The first mechanism changes the order of release. Today all reviews and ratings reach authors at once. Under the proposal, each review's summary, strengths and clarifying questions come first, and authors grade them on the reviewer's comprehension and the constructiveness of the questions. Weaknesses and ratings follow, and authors cannot revise their evaluations afterward. The authors argue that rating a review before seeing its verdict blocks "retaliatory scoring." An LLM review, visible only to authors, is meant to be "a psychological deterrent" for reviewers and "a soft reference point" for spotting machine-written reviews. For rewards, reviewers would earn verifiable badges. The paper concedes that badges could reward lenient reviewing, and asks for a metric that rewards thoroughness over positivity.
The diagnosis rests on NeurIPS studies, as the paper reports them. In consistency experiments in 2014 and 2021, "16-23% of papers could have been either accepted or rejected based on the reviewing group." A 2022 NeurIPS study of review quality found "author-outcome bias" and "elongated review bias," the latter rating longer reviews higher "even when the two reviews contain the same information." The two-stage design targets the first, and capping the length of stage-one content targets the second. Against the nearest notes, this is a different lever from the one the ICML 2026 trial tested. Does banning LLM use in peer review change review outcomes? found rules barely moved outcomes, while this paper bets on visibility, which it does not test. Its LLM review is a reference for authors, not a reviewer, unlike Can inference scaling help reviewers catch errors humans miss?, where a model checks proofs and experiments line by line. The length bias also parallels the surface sensitivity in How much does rhetorical style shift AI review scores?, here measured in human reviewers.
The excerpt does not establish that the fix works. The paper reports no pilot. Its Section 5.1 asks for a large survey and a small-track pilot first, and Section 5.2 names the main obstacle as getting venues and OpenReview to build the feedback and reward systems, plus cost: ICML 2024 cut in-kind compensation for its top 10% of reviewers to a smaller pool. The claim that stage-one release "could prevent retaliatory scoring" is an inference from the bias finding, not a measured effect, and the NeurIPS results come without effect sizes. The 10,000-submissions figure carries no citation in the excerpt. The strongest support is for the diagnosis, that review is noisy and rating-biased. The prescription is a plausible design that a venue could test as a pilot before adopting it as policy.
Inquiring lines that read this note 46
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans?- Can multi-stage AI review pipelines catch scientific flaws better than simple language models?
- Did adding AI reviews actually change peer review decisions or paper outcomes?
- Should rhetorical polish in AI reviews be separated from actual technical accuracy?
- How much of ICLR 2026 peer review was already conducted by AI?
- Do AI reviews depend more on writing style than scientific merit?
- Should AI research papers require dedicated automated review systems instead of existing journals?
- Can workshop acceptance rates reliably measure AI research quality compared to main conferences?
- Does review length bias affect acceptance decisions at major conferences?
- How much does reviewer consistency vary across different papers at NeurIPS?
- Could AI improve peer review rigor and catch human-missed errors?
- Does matching reviewer points actually mean the feedback is accurate or correct?
- Could AI feedback work as a substitute for human peer review entirely?
- Can computational inference scaling catch flaws that human expert reviewers miss?
- What effects do preprint servers have on scientific consensus formation?
- Can institutional statements alone correct misconceptions from unreviewed papers?
- Do AI-generated research reviews score papers higher than human reviewers do?
- How often do researchers suspect peer reviews are written by AI?
- Does AI content in reviews correlate with differences in paper quality control?
- Can human reviewers reliably detect AI-written peer review text by sight?
- How do automated reviewers detect flaws that human experts miss in manuscripts?
- Should AI-generated papers use specialized review venues instead of traditional journals?
- Why do peer reviewers favor novel ideas that later fail in execution?
- Do peer reviewers actually follow restrictions on using AI tools themselves?
- Are refereed venues also overwhelmed by AI-generated low-quality submissions?
- Can automated systems scale peer review faster than human moderators?
- Does rhetorical presentation bias reviewers against substantive scientific contributions?
- Can agentic AI systems catch flaws in manuscripts that human reviewers consistently miss?
- Why do individual peer reviewers show such low agreement on research merit?
- How fast is scientific publishing growing relative to reviewer capacity?
- Could automated review systems handle AI-generated research at scale?
- What role do conference organizers play in accepting problematic articles?
- Which feedback loops in AI-mediated review remain unmeasured or rarely observed directly?
- Do shortened peer review timelines correlate with lower quality publications?
- How do AI-generated papers perform when submitted to real conferences?
- Why do researchers resist using AI for peer review specifically?
- Do academic reward structures actively prevent innovation in research communication forms?
- Can automated reviewers actually handle the review load AI creates?
- Can traditional complexity measures still signal research quality in AI-era papers?
- What prevents venues from implementing two-way feedback systems at scale?
- Do reviewer reward badges risk encouraging lenient or superficial reviews?
- Do stricter AI policies actually change how reviewers score manuscripts?
- Do conference policies banning LLM use actually reduce AI involvement in reviews?
- Why do peer review policies often fail to change actual review scores?
- What happens when reviewers use AI tools against journal policy?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
contrasts: rule-based LLM policy that barely moved outcomes; this paper proposes visibility instead, untested.
-
Can automated review loops handle AI-generated research at scale?
As AI agents produce papers faster than humans can evaluate them, can a closed-loop automated review system with retrieval-augmented feedback actually improve quality and catch problems traditional peer review misses?
parallel review loop, but machine-run for AI papers; this paper's loop runs between people.
-
Can inference scaling help reviewers catch errors humans miss?
Explores whether spending extra compute at review time—checking proofs and experiments line by line—can surface deep flaws that evade human expert reviewers, and how this scales with AI-assisted submissions.
contrasts: an LLM as the reviewer; this paper uses LLM output only as an author reference.
-
How much does rhetorical style shift AI review scores?
When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.
parallel: surface features move scores in LLM reviewers, and the length bias shows it in humans.
-
Can LLM feedback help peer reviewers improve their own reviews?
This randomized trial tested whether optional AI-generated suggestions on review quality would prompt reviewers to revise, and whether those revisions would be more useful to authors and decision-makers.
evidence for: a randomized ICLR 2025 trial found optional LLM comments on reviews led about a quarter of recipients to revise them
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Stop Automating Peer Review Without Rigorous Evaluation
- How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
Original note title
AI conference review problems are shared by three parties, so the position argues for two-way author feedback and reviewer rewards