Give an AI reviewer more time to check line by line, and it can catch real errors top conferences let through.
Can computational inference scaling catch flaws that human expert reviewers miss?
This explores whether giving an AI reviewer more time to think while it works (more checking steps, more reasoning, more evidence gathering) lets it find real errors in research papers that expert human reviewers missed, and what that does and doesn't mean for peer review.
This explores whether giving an AI reviewer more thinking time and more checking steps lets it find real errors in papers that expert human reviewers let through. The corpus has one direct answer, and it is a qualified yes. An agentic reviewer called PAT spends extra inference compute walking through proofs and experiments line by line. It found 34% more mathematical errors than a model asked to review in a single pass, and it found critical flaws in papers already accepted at STOC (a top theoretical computer science conference) and ICML (a major machine learning conference) Can inference scaling help reviewers catch errors humans miss?. Human reviewers don't miss these errors because they can't understand them. They miss them because line-by-line checking is tedious, and few people have time to rework every step of a 40-page appendix. Extra compute is good at exactly that kind of patient work.
The word doing the work there is *agentic*. More raw thinking time is not the main source of the gain. The gain comes from a structure that gathers evidence and checks it. An agent-based evaluator that actively collects evidence was about 100 times more stable in its verdicts than a plain LLM judge Can agents evaluate AI outputs more reliably than language models?. A paper-writing system improves reliability by separating the model's judgment from steps that can be run and checked mechanically Can separating judgment from verification improve research paper reliability?. The same pattern shows up in reasoning research. Models that 'collapse' on long problems often know the right method but can't carry out hundreds of steps as plain text, and tools let them push past that limit Are reasoning model collapses really failures of reasoning?. Compute also only pays off when the model was trained to use it well: non-reasoning models don't catch up to reasoning models however many tokens they get Can non-reasoning models catch up with more compute?. Search steps follow a similar curve with diminishing returns Do search steps follow the same scaling rules as reasoning tokens?.
Here is the twist. Catching flaws humans miss is not the same as being a good reviewer. AI reviewers show a 'hivemind' effect: they agree with each other far more than human reviewers do, so their blind spots are shared rather than offset. Simply rewording a paper, with no change to the science, raised AI scores by almost half a point Can AI systems safely replace human peer reviewers?. Reasoning models also tend to fail on unfamiliar *instances* rather than on hard ones Do language models fail at reasoning due to complexity or novelty?. That is a worry for review, because the papers that matter most are the novel ones. So machines and humans fail in different places. Machines are strong at exhaustive checking and weak at judging novelty. Humans are the reverse, and that same gap let an AI-generated paper pass a workshop's blind review before its own authors later found a citation error in it Can AI-generated papers pass peer review undetected?.
That points to the most promising setup: machine checking that supports human reviewers rather than replacing them. In a randomized trial at ICLR 2025, optional AI feedback on reviews led 27% of reviewers to revise, and independent raters judged the revised reviews more informative Can LLM feedback help peer reviewers improve their own reviews?. Proposals to repair conference review already treat it as a shared-accountability problem between authors, reviewers and venues Can two-stage review and badges fix AI conference peer review?. An inference-scaled checker fits naturally as a fourth party that does the tedious verification. The surprise is which flaws inference scaling catches. They aren't subtle insights beyond human reach. They are mechanical errors that humans could find but rarely have time to look for.
Sources 11 notes
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.
Models confined to text-only generation cannot execute multi-step procedures at scale, even when they know the underlying algorithm. Tool-enabled models solve problems beyond the supposed reasoning cliff, suggesting the bottleneck is procedural execution bandwidth.
Reasoning models persistently outperform non-reasoning models regardless of inference budget because training instills a reasoning protocol that makes additional tokens productive. The gap is fundamentally about deployment mechanisms and training structure, not raw capability.
Show all 11 sources
Deep research agents improve with more search steps in a pattern mirroring the reasoning-token relationship, with both exhibiting diminishing returns. This reveals a new inference-compute axis beyond model capability alone.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
LRMs don't break at complexity thresholds but at instance-novelty boundaries. Models fit instance-based patterns rather than generalizable algorithms, so any reasoning chain succeeds if trained on similar instances, regardless of length.
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing