INQUIRING LINE

Give an AI reviewer more time to check line by line, and it can catch real errors top conferences let through.

Can computational inference scaling catch flaws that human expert reviewers miss?

This explores whether giving an AI reviewer more time to think while it works (more checking steps, more reasoning, more evidence gathering) lets it find real errors in research papers that expert human reviewers missed, and what that does and doesn't mean for peer review.


This explores whether giving an AI reviewer more thinking time and more checking steps lets it find real errors in papers that expert human reviewers let through. The corpus has one direct answer, and it is a qualified yes. An agentic reviewer called PAT spends extra inference compute walking through proofs and experiments line by line. It found 34% more mathematical errors than a model asked to review in a single pass, and it found critical flaws in papers already accepted at STOC (a top theoretical computer science conference) and ICML (a major machine learning conference) Can inference scaling help reviewers catch errors humans miss?. Human reviewers don't miss these errors because they can't understand them. They miss them because line-by-line checking is tedious, and few people have time to rework every step of a 40-page appendix. Extra compute is good at exactly that kind of patient work.

The word doing the work there is *agentic*. More raw thinking time is not the main source of the gain. The gain comes from a structure that gathers evidence and checks it. An agent-based evaluator that actively collects evidence was about 100 times more stable in its verdicts than a plain LLM judge Can agents evaluate AI outputs more reliably than language models?. A paper-writing system improves reliability by separating the model's judgment from steps that can be run and checked mechanically Can separating judgment from verification improve research paper reliability?. The same pattern shows up in reasoning research. Models that 'collapse' on long problems often know the right method but can't carry out hundreds of steps as plain text, and tools let them push past that limit Are reasoning model collapses really failures of reasoning?. Compute also only pays off when the model was trained to use it well: non-reasoning models don't catch up to reasoning models however many tokens they get Can non-reasoning models catch up with more compute?. Search steps follow a similar curve with diminishing returns Do search steps follow the same scaling rules as reasoning tokens?.

Here is the twist. Catching flaws humans miss is not the same as being a good reviewer. AI reviewers show a 'hivemind' effect: they agree with each other far more than human reviewers do, so their blind spots are shared rather than offset. Simply rewording a paper, with no change to the science, raised AI scores by almost half a point Can AI systems safely replace human peer reviewers?. Reasoning models also tend to fail on unfamiliar *instances* rather than on hard ones Do language models fail at reasoning due to complexity or novelty?. That is a worry for review, because the papers that matter most are the novel ones. So machines and humans fail in different places. Machines are strong at exhaustive checking and weak at judging novelty. Humans are the reverse, and that same gap let an AI-generated paper pass a workshop's blind review before its own authors later found a citation error in it Can AI-generated papers pass peer review undetected?.

That points to the most promising setup: machine checking that supports human reviewers rather than replacing them. In a randomized trial at ICLR 2025, optional AI feedback on reviews led 27% of reviewers to revise, and independent raters judged the revised reviews more informative Can LLM feedback help peer reviewers improve their own reviews?. Proposals to repair conference review already treat it as a shared-accountability problem between authors, reviewers and venues Can two-stage review and badges fix AI conference peer review?. An inference-scaled checker fits naturally as a fourth party that does the tedious verification. The surprise is which flaws inference scaling catches. They aren't subtle insights beyond human reach. They are mechanical errors that humans could find but rarely have time to look for.


Sources 11 notes

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Are reasoning model collapses really failures of reasoning?

Models confined to text-only generation cannot execute multi-step procedures at scale, even when they know the underlying algorithm. Tool-enabled models solve problems beyond the supposed reasoning cliff, suggesting the bottleneck is procedural execution bandwidth.

Can non-reasoning models catch up with more compute?

Reasoning models persistently outperform non-reasoning models regardless of inference budget because training instills a reasoning protocol that makes additional tokens productive. The gap is fundamentally about deployment mechanisms and training structure, not raw capability.

Show all 11 sources
Do search steps follow the same scaling rules as reasoning tokens?

Deep research agents improve with more search steps in a pattern mirroring the reasoning-token relationship, with both exhibiting diminishing returns. This reveals a new inference-compute axis beyond model capability alone.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Do language models fail at reasoning due to complexity or novelty?

LRMs don't break at complexity thresholds but at instance-novelty boundaries. Models fit instance-based patterns rather than generalizable algorithms, so any reasoning chain succeeds if trained on similar instances, regardless of length.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.