INQUIRING LINE

In one study, an AI matched human reviewers' novelty verdicts about three-quarters of the time, but only once it was made to reason like one.

Do humans and LLMs agree on novelty assessment in research?

This explores whether LLMs judge how new a research idea or paper is the same way human reviewers do, and what the collection says about where their judgments line up and where they part ways.


This explores whether LLMs and human experts reach the same verdicts when judging how new a piece of research is. The short answer from the collection: they can agree surprisingly well, but only when the LLM is made to work the way a careful reviewer does. A three-step pipeline (pull out the paper's claims, find related work, then compare the two) matched human reviewers' reasoning 86.5% of the time and their final conclusion about three-quarters of the time on ICLR submissions. It also clearly beat simply asking a model whether a paper is novel Can structured pipelines make LLM novelty assessment reliable?. So agreement isn't something the model has by default. It comes from the process, and the comparison against prior work does most of the work.

Things get stranger once you flip from judging novelty to producing it. In a large study with 100+ NLP researchers, human experts rated LLM-generated research ideas as more novel than ideas written by other experts Do language models generate more novel research ideas than experts?. One reading is that LLMs combine concepts freely because they lack the disciplinary instincts that tell experts what's off-limits. Yet the same models tend to avoid taking the evaluative stance needed to say whether an idea would actually work Can LLMs generate more novel ideas than human experts?. Generating something new and judging something new turn out to be separate skills, and a model can be strong at one and weak at the other.

Here is the twist that changes the question: humans don't agree with themselves over time. When 43 researchers spent 100+ hours actually carrying out randomly assigned ideas, the LLM ideas that had looked more novel lost far more quality than the human ideas did. Execution exposed impractical evaluation plans and missing technical groundwork Do LLM research ideas actually hold up when experts try to execute them?. So the 'humans' in 'do humans and LLMs agree' are partly reacting to how an idea reads on the page. Novelty judged at the proposal stage and novelty judged after the work is done can be very different. And it doesn't hold across fields: in engineering design, experts rated LLM concepts as more feasible but less novel than crowdsourced human ones, and few-shot prompting narrowed the variety further Why do LLMs excel at feasible design but struggle with novelty?.

There are also reasons LLM judges can drift from human ones on purpose-built grounds. Models tend to prefer text they recognize as their own Do LLMs favor their own text because they recognize it?, which matters when they're judging work that more and more often passes through an LLM. That includes abstracts that readers already can't reliably tell apart from human writing Can readers tell LLM abstracts from human ones?. And models only see text, without the reputation and track record that tell a human reviewer whether a claim comes from someone who has earned the right to make it Can language models distinguish expert arguments from common assumptions?. Conferences seem to have settled on keeping humans in charge. ICLR used LLM feedback to sharpen human reviews rather than replace them, and 27% of reviewers revised Can LLM feedback help peer reviewers improve their own reviews?. It also treated LLM detectors as one signal for area chairs rather than an automatic filter How can conferences detect and handle LLM misuse in peer review?.

What you might not expect: where novelty can be checked automatically against a hard test, the agreement question mostly goes away. Systems like AlphaEvolve don't need anyone's opinion about whether a faster algorithm is new and better, because the evaluator can simply measure it Can machine feedback sustain discovery at test time?. Human-LLM disagreement about novelty is mostly a problem in fields where 'new' can't be verified cheaply. The collection doesn't yet have a head-to-head study of LLM and human novelty scores on the same papers checked against what happened to those papers later, and that is the missing piece.


Sources 11 notes

Can structured pipelines make LLM novelty assessment reliable?

A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.

Do language models generate more novel research ideas than experts?

A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.

Can LLMs generate more novel ideas than human experts?

LLMs produce more novel research ideas than experts because they lack disciplinary constraints, but they systematically avoid evaluative stance-taking required to assess feasibility or validity. Generation and evaluation are dissociated capabilities.

Do LLM research ideas actually hold up when experts try to execute them?

When 43 expert researchers implemented randomly-assigned ideas over 100+ hours, LLM-generated ideas declined significantly more than human ideas across all metrics. Execution revealed systematic weaknesses invisible at ideation, including impractical evaluation designs and missing technical groundwork.

Why do LLMs excel at feasible design but struggle with novelty?

Expert evaluation shows LLM-generated conceptual designs score higher on feasibility and usefulness but lower on novelty compared to crowdsourced human solutions. Few-shot learning further reduces diversity while improving quality alignment.

Show all 11 sources
Do LLMs favor their own text because they recognize it?

Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Can language models distinguish expert arguments from common assumptions?

LLMs lose the social context that gives expert claims their force—reputation, track record, and standing—because they process only text, not the social world where expertise is built and evaluated.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Can machine feedback sustain discovery at test time?

AlphaEvolve demonstrates that automated evaluators can sustain evolutionary loops long enough to produce real discoveries—faster algorithms, optimized hardware designs, and improved training methods. The key is that cheap, objective verification closes the generation-verification gap where discovery becomes computationally feasible.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.