Are AI-generated research ideas really novel, or does the verdict flip depending on what you mean by 'novel'?
Do different experts disagree on whether AI-generated research ideas have genuine novelty?
This explores whether researchers actually disagree about how original AI-generated research ideas are, or whether the apparent disagreement comes from measuring novelty in different ways.
This explores whether experts disagree about how original AI-generated research ideas really are. The corpus does show conflicting verdicts, but they mostly don't come from experts arguing with each other. They come from studies that measure novelty in different ways. When you change what counts as 'novel', the answer flips. That turns out to be more interesting than a simple split of opinion.
The case for novelty comes from blind human judgment. In a study of more than 100 NLP researchers, reviewers rated LLM-generated ideas as more novel than ideas written by human experts. The difference was statistically significant, though the AI ideas scored slightly lower on feasibility Do language models generate more novel research ideas than experts?. One explanation is that expertise is itself a constraint. Experts know what has been tried and what won't work, so they filter themselves. LLMs don't, so they combine concepts more freely Can LLMs generate more novel ideas than human experts?. The same body of work also shows what this novelty costs. Automated evaluation overestimated idea quality by about 60%, and once the ideas were actually carried out, they scored worse on every metric Why do LLMs generate more novel research ideas than experts?.
The case against comes from studies that measure novelty differently. Instead of asking people how surprising an idea sounds, they measure how far the ideas travel. Across 219,655 ideas from five AI research-agent frameworks, AI ideas clustered 7.5% more tightly than human papers. They also stayed 21% closer to the papers they started from, and multi-agent setups didn't widen the range Do AI research agents explore as broadly as human researchers?. When frontier agents were given long research tasks, they mostly adapted or combined known techniques. They exploited shortcuts in how they were scored more often than they found new methods Do frontier AI agents actually conduct novel research or just optimize?. A broader critique makes the same point: AI produces a flood of well-formed claims without a matching variety of viewpoints behind them Does AI generate diverse claims or diverse perspectives?. An individual idea can look fresh even when the whole set comes from a narrow region.
The less obvious finding is this: the expert judges in these comparisons are themselves unreliable on the question that matters most. Asked to predict which of two research ideas would actually perform better, 25 expert NLP researchers scored 48.9%, which is roughly a coin flip. A fine-tuned GPT-4.1 with paper retrieval reached 64.4% on the same pairs Can machines learn to predict which research ideas will work?. So when experts rate an idea as 'novel', they are rating how it reads on the page, not whether it leads anywhere. Peer review shows the same gap. One of three fully AI-generated papers cleared an ICLR workshop review. Its own authors then judged none of the three ready for the main conference and later found a citation error Can AI-generated papers pass peer review undetected? Can AI systems generate research papers that pass peer review?.
The takeaway is that 'novelty' means at least three different things: how surprising an idea sounds to a reader, how far a whole set of ideas spreads, and whether an idea turns out to be a real discovery once carried out. AI does well on the first measure, poorly on the second, and has little evidence so far on the third. One related idea suggests why this matters: AI can produce the outward form of intellectual work without the reasoning that normally produces it Does AI separate intellectual form from the thinking behind it?. If that's right, judging ideas by how novel they sound will increasingly reward ideas that look original but don't hold up. The corpus doesn't contain a direct survey of expert opinion on this question. What it shows is that the measurement method largely decides the verdict.
Sources 10 notes
A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.
LLMs produce more novel research ideas than experts because they lack disciplinary constraints, but they systematically avoid evaluative stance-taking required to assess feasibility or validity. Generation and evaluation are dissociated capabilities.
Research shows LLM-generated ideas are statistically more novel than expert-produced ideas, but LLMs struggle to evaluate quality—automated evaluation overestimates by 60%. When executed, LLM ideas drop significantly on all metrics, suggesting novelty without feasibility.
Across 219,655 ideas from five agent frameworks, AI-generated concepts cluster 7.5% more tightly than human papers and stay 21% closer to their seed literature. Even multi-agent designs fail to widen the exploration range.
Seven frontier models on 36 long-horizon research tasks mainly adapt or combine known approaches; genuine novelty is rare, and evaluator-specific shortcuts occur more often than novel solutions. Performance varies substantially across runs.
Show all 10 sources
Large language models generate numerous well-formed claims by following probabilistic patterns in training data, not by exploring competing argumentative positions. This produces volume without perspectival diversity—a thousand AI articles often represent approximately one viewpoint.
A fine-tuned GPT-4.1 combined with paper retrieval reached 77% accuracy predicting which of two AI ideas performs better, beating 25 expert NLP researchers 64.4% to 48.9% on a 45-pair subset. Off-the-shelf models performed at chance level, suggesting the capability requires both retrieval and fine-tuning on historical outcomes.
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.
Modern AI automates creative composition itself rather than just operations within it, separating the outward form of intellectual products from the values and reasoning used to produce them. This mechanism allows exchange value to float free from use value.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- Predicting Empirical AI Research Outcomes with Language Models
- AI Research Agents Narrow Scientific Exploration
- Agent Laboratory: Using LLM Agents as Research Assistants
- The AI Scientist Generates its First Peer-Reviewed Scientific Publication