INQUIRING LINE

Why might a paper's exciting idea win over reviewers, while the hard part, making it work, goes unchecked?

Why do peer reviewers favor novel ideas that later fail in execution?

This explores why peer review might reward ideas that sound new and exciting on paper but don't hold up when someone actually carries them out. In other words, what peer review can and can't see when it judges a submission.


This explores why peer review might reward ideas that sound new and exciting on paper but don't hold up when someone actually carries them out. One caveat first: none of the retrieved notes directly tests the claim that reviewers prefer novel ideas that later fail. What the collection does have is evidence about the conditions that would produce that pattern. Review rewards what a reviewer can judge quickly from the page, and the hard parts of execution often can't be checked in the time a reviewer has.

Start with what reviewers can't see. An agentic reviewer that spends extra compute checking proofs and experiments line by line found critical flaws in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. Human reviewers aren't careless. Execution errors are buried in the details, and reviewers judge the overall story. Novelty is part of that story and is easy to recognize. Whether the method actually works is not. Novelty itself can be broken into steps: pull out the paper's claims, find the related work, and compare. A pipeline built that way matched human reviewers' reasoning 86 percent of the time Can structured pipelines make LLM novelty assessment reliable?. Novelty is a checkable judgment, and review is built around making it.

The pressure makes this worse. One model shows a feedback loop: more submissions overload unpaid reviewers, so journals recruit less qualified ones, accuracy drops, and authors respond by submitting more speculative work Does peer review quality collapse under submission overload?. When reviewers are stretched, surface signals carry more weight. One measured bias is that review scores track how long a review is Can two-stage review and badges fix AI conference peer review?. AI reviewers are even easier to sway: rewording a paper with no change to its science raised their scores by almost half a point Can AI systems safely replace human peer reviewers?. A fully AI-generated paper scored above the acceptance threshold at an ICLR workshop, and its own authors later found a citation error and judged none of their three submissions ready for the main conference Can AI-generated papers pass peer review undetected?. A paper can look convincing without being solid.

Here is the result you might not expect. At ICML 2023, authors' private rankings of their own submissions predicted citations over the following 16 months better than the official review scores did Can authors rank their own papers better than peer reviewers?. Authors know things reviewers don't: which experiments were fragile, which results they trust, which idea they could actually deliver. That points to a cause beyond reviewers being drawn to shiny ideas. Review is a one-shot reading by someone with less information than the author, so it can't price in the risk of execution. One proposed fix is two-stage review, where authors rate the quality of the reviews before seeing the decision. Proposals like this try to bring that hidden author knowledge back into the process Can two-stage review and badges fix AI conference peer review?.

If you want the head-to-head evidence, meaning studies that track reviewed ideas through execution and compare predicted and actual outcomes, the collection doesn't have it here. The closest doorways are the self-ranking study and the overload model above.


Sources 7 notes

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can structured pipelines make LLM novelty assessment reliable?

A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.

Does peer review quality collapse under submission overload?

A two-journal model shows that rising submissions overtax unpaid reviewers, forcing journals to recruit less qualified reviewers or overload existing ones, which drops review accuracy and incentivizes authors to submit more speculatively, driving submissions higher. The mechanism is structural but its empirical strength remains to be measured.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Show all 7 sources
Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can authors rank their own papers better than peer reviewers?

At ICML 2023, self-rankings by 1,342 researchers predicted future citations better than peer review scores over 16 months. Top-ranked papers drew twice the citations of bottom-ranked ones, and 77% of highly-cited papers had been ranked highest by their authors.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.