When GPT-4 reviews a paper, its points line up with one human reviewer about as well as two reviewers' points line up with each other.
How did researchers measure whether GPT-4 and human reviewers identified the same issues?
This explores how researchers tested whether GPT-4's comments on a scientific paper hit the same points human reviewers raised, and what that kind of measurement can and can't tell us.
This explores how researchers checked whether GPT-4 and human peer reviewers flagged the same problems in a paper. The approach did not ask whether GPT-4 was 'right.' It asked a narrower question: when GPT-4 reviews a paper, how many of its points line up with points a human reviewer made on that same paper? The key design choice was to use humans as the baseline. Across 3,096 Nature-family papers and 1,709 ICLR papers, GPT-4's comments matched an individual reviewer's points 30.85% of the time. Two human reviewers of the same paper matched each other 28.58% of the time Can GPT-4 feedback match what human reviewers catch?. That comparison is what makes the number readable. On its own, 30% overlap sounds low. But human reviewers also disagree most of the time about what matters, so GPT-4 lands inside the normal range of variation between reviewers. The researchers added a second, separate measure: a survey in which 57% of researchers said they found the feedback helpful. Note that the corpus note gives the results, not the step-by-step method for deciding when two comments count as 'the same point.'
The surprising part is what overlap misses. Another line of work finds a 'hivemind effect.' AI reviewers agree with each other more than humans do, across many papers. And simply rewriting a paper's wording, with no change to the science, raises AI scores by about 0.45 points Can AI systems safely replace human peer reviewers?. So a model can overlap with humans at a human-like rate while still lacking the variety of viewpoints that makes having several reviewers worthwhile. A single 'does it match?' number can't show that. You have to compare AI reviewers with each other too. This is also why some evaluation research prefers panels of smaller models from different families over one large judge. Diversity among judges reduces shared blind spots Can a panel of smaller judges outperform one large judge?.
Other studies in the collection measure 'did the AI catch what humans catch' in quite different ways, and the contrast is useful. The agentic reviewer PAT is scored on recall: how many known mathematical errors it finds. It even surfaced flaws in STOC and ICML papers that had passed human review Can inference scaling help reviewers catch errors humans miss?. There, humans are no longer the standard. They are a group the AI might beat. The ICLR 2025 trial measured something else again: whether LLM feedback changed what human reviewers wrote. Blinded raters then judged the revised reviews (27% of reviewers updated theirs) as more informative Can LLM feedback help peer reviewers improve their own reviews?.
Outside peer review, blinded comparison is a common way to measure 'same as a human,' and it carries a warning. Clinicians couldn't tell GPT-4's medical advice from expert advice (45% accuracy, roughly chance) Can clinicians tell GPT-4 advice apart from expert advice?. Evaluators have also been fooled by models that copy ChatGPT's confident style without matching its accuracy Can imitating ChatGPT fool evaluators into thinking models improved?. The takeaway for reading the GPT-4 reviewer result is that matching humans and being as good as humans are different claims. Overlap tells you GPT-4 comments on similar things. It doesn't tell you whether those comments are correct, independent of other models, or hard to game.
Sources 7 notes
A study of 3,096 Nature papers and 1,709 ICLR papers found GPT-4 matched individual reviewers' points 30.85% of the time versus 28.58% for two human reviewers. Fifty-seven percent of surveyed researchers found the feedback helpful.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
PoLL (Panel of LLM evaluators) using multiple smaller models from disjoint families outperforms single large judges, reduces intra-model bias, and costs over 7× less. Across three settings and six datasets, no single judge was best everywhere, but diverse panels performed consistently well.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Show all 7 sources
Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.
Imitation models fool human evaluators by mimicking ChatGPT's confident, fluent style while failing to improve factuality or generalization on novel tasks. The ceiling is set by base model capability, not fine-tuning method—better fundamentals, not shortcuts, drive real improvement.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review