INQUIRING LINE

When top AI conference reviewers score the same paper, how often do they actually agree, and does that vary by paper?

How much does reviewer consistency vary across different papers at NeurIPS?

This explores how much reviewers agree or disagree when scoring the same papers at a big AI conference like NeurIPS, and whether that agreement shifts from paper to paper.


This explores how much reviewers agree with each other when they score the same NeurIPS papers, and whether some papers draw much more disagreement than others. To be direct: this collection doesn't hold the classic NeurIPS consistency experiments, where the same submissions went to two independent committees and the accept/reject decisions were compared. So it can't give you a number for how much consistency varies. What it does have is a set of nearby findings. Together they explain why reviewer agreement is hard to pin down, and why making reviewers agree more isn't automatically better.

The most surprising angle comes from AI reviewers. When researchers compared AI and human reviews across many papers, the AI systems agreed with each other far more than human reviewers did. That is a 'hivemind effect,' and the researchers count it as a failure, not a strength Can AI systems safely replace human peer reviewers?. Some of the disagreement between human reviewers is the point of having several of them, because it brings different expert views to the same paper. The same AI reviewers could also be gamed: simply rewording a paper raised its scores by almost half a point with no change to the science. Higher agreement can mean everyone is reacting to the same surface signal.

Other notes look at what shapes human scores apart from paper quality. One position paper on AI conference review points to measured biases, including a link between how long a review is and the rating it gives. It argues that authors, reviewers and venues all share responsibility for these problems Can two-stage review and badges fix AI conference peer review?. A study of more than 125,000 reviews found an apparent pattern that dissolved under scrutiny. LLM-assisted reviewers seemed to favor LLM-written papers, but the effect disappeared once paper quality was held constant. The real pattern was that these reviewers were more lenient toward weaker papers in general Do LLM reviewers actually favor LLM-written papers?. That kind of reviewer-specific leniency is exactly what makes scores swing from paper to paper depending on who happens to be assigned.

Two randomized trials test whether changing the review process shifts outcomes. At ICLR 2025, optional AI feedback on draft reviews led 27% of reviewers to revise, and blinded raters judged the revisions more specific Can LLM feedback help peer reviewers improve their own reviews?. At ICML 2026, banning LLM use versus allowing limited use barely moved scores or decisions. Many reviewers broke whichever rule they were given anyway Does banning LLM use in peer review change review outcomes?. One more idea comes from outside academia. Research on online product ratings shows that early ratings pull later ones toward them, so a rating is never fully independent Do online ratings actually reflect independent customer opinions?. If the same thing happens in reviewer discussions, apparent agreement on a paper may partly reflect reviewers moving toward whoever spoke first.

If what you want is NeurIPS specifically, the collection mostly looks at the quality of accepted papers rather than reviewer agreement. For example, a scan of 4,841 accepted NeurIPS 2025 papers flagged hundreds of possibly hallucinated citations that got past review How many accepted conference papers contain hallucinated citations?. That's a different kind of reliability gap: reviewers may agree on a verdict and still all miss the same thing.


Sources 7 notes

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Do LLM reviewers actually favor LLM-written papers?

Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

Show all 7 sources
Do online ratings actually reflect independent customer opinions?

Moe and Trusov decomposed ratings into baseline quality, social-dynamics influence, and error, finding that prior ratings meaningfully affect subsequent ones. These effects have both immediate sales impact and long-term compounding effects through future ratings, though high opinion variance can eventually dampen the distortion.

How many accepted conference papers contain hallucinated citations?

GPTZero's citation checker flagged hundreds of potentially hallucinated citations across 4841 accepted NeurIPS 2025 papers. However, flagged citations require human verification to confirm hallucination, and the full verification rate across the full scan remains undisclosed.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.