INQUIRING LINE

Ask an AI for novel ideas and each one can look fresh, but across many runs they bunch into a few narrow corners.

Why do AI agents pursue novelty prompts yet produce narrow idea ranges?

This explores a puzzle: when you ask AI agents to be novel, each idea can look fresh on its own, yet across many runs the ideas pile up in a small corner of the possible space. Why do both things happen at once?


This explores why AI agents asked for novelty give you ideas that each seem original but, taken together, cover surprisingly little ground. The short answer from the corpus: novelty and diversity are different things, and the model is good at the first and bad at the second. In a large study with more than 100 NLP researchers, experts rated LLM-generated research ideas as more novel than human ones, though a bit less feasible Do language models generate more novel research ideas than experts?. But when you look at the whole set of ideas rather than one at a time, they collapse into narrow clusters Why do LLMs generate novel ideas from narrow ranges?. A grader reading one idea at a time sees something surprising. A map of all the ideas shows the same few neighborhoods visited again and again.

The largest measurement in the collection makes this concrete. Across 219,655 ideas from five agent frameworks, AI ideas sat 7.5% more tightly clustered than human papers and stayed 21% closer to the literature they started from Do AI research agents explore as broadly as human researchers?. That fits what happens when agents actually do the research. On 36 long-horizon tasks, seven frontier models mostly adapted or recombined known techniques. They found shortcuts that exploited the evaluator more often than they found new methods Do frontier AI agents actually conduct novel research or just optimize?. So what reads as "novel" is usually an unusual combination of familiar parts. The parts are new to each other, but each one comes from close by.

Why close by? Two mechanisms show up in other parts of the corpus under different names. First, training sets a ceiling. Agents trained on expert demonstrations can't go beyond what the people who built the data imagined Can agents learn beyond what their training data shows?. Second, reinforcement learning actively narrows behavior. In search agents, RL squeezes exploration down to a few strategies that maximize reward, the same "entropy collapse" seen in reasoning models, while fine-tuning on diverse examples keeps the range wider Does reinforcement learning squeeze exploration diversity in search agents?. A prompt that says "be novel" can't undo either one. The model reaches for the most surprising-sounding option within the space it already favors.

The less obvious finding is that adding more agents doesn't solve this. The 219K-idea study found that multi-agent designs did not widen the range Do AI research agents explore as broadly as human researchers?. A separate line of work helps explain why: a single LLM playing several personas reproduces what multi-agent debate does Can branching prompts replicate what multi-agent systems do?. If that works, it also means several copies of one model are mostly drawing from the same distribution. Their "different perspectives" are variations on one underlying view. Diversity among agents only helps when each one brings real domain expertise. Teams that are diverse without that expertise do worse than a single competent agent Does cognitive diversity alone improve multi-agent ideation quality?.

If you're looking for what actually widens the range, the corpus points to structure rather than exhortation. One approach first generates several different high-level abstractions (ways of framing the problem) and then solves within each. That forces breadth-first exploration, and at large compute budgets it beats sampling many solutions in parallel Can abstractions guide exploration better than depth alone?. Telling a model to be original changes how its ideas sound. Changing where it starts changes how far apart its ideas end up. The corpus doesn't have a study that tests novelty prompts directly against these structural fixes, so that comparison is still open.


Sources 9 notes

Do language models generate more novel research ideas than experts?

A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.

Why do LLMs generate novel ideas from narrow ranges?

LLM-generated research ideas are rated individually novel but lack diversity, clustering in narrow generative regions. Combined with LLM self-evaluation failures, this limits the possibility space explored compared to human ideation across different conceptual territories.

Do AI research agents explore as broadly as human researchers?

Across 219,655 ideas from five agent frameworks, AI-generated concepts cluster 7.5% more tightly than human papers and stay 21% closer to their seed literature. Even multi-agent designs fail to widen the exploration range.

Do frontier AI agents actually conduct novel research or just optimize?

Seven frontier models on 36 long-horizon research tasks mainly adapt or combine known approaches; genuine novelty is rare, and evaluator-specific shortcuts occur more often than novel solutions. Performance varies substantially across runs.

Can agents learn beyond what their training data shows?

Agents trained on static expert datasets cannot learn from their own failures or generalize beyond demonstrated scenarios because they never interact with environments during training. Competence is capped by what curators imagined, not by agent capacity.

Show all 9 sources
Does reinforcement learning squeeze exploration diversity in search agents?

RL training compresses behavioral diversity in search agents through the same entropy collapse mechanism documented in reasoning—policies converge on narrow reward-maximizing strategies. SFT on diverse demonstrations preserves exploration breadth, suggesting diversity-preservation techniques are essential for RL search scaling.

Can branching prompts replicate what multi-agent systems do?

Research shows single LLMs using dynamic persona simulation achieve multi-agent cognitive synergy without multiple model instances. Solo Performance Prompting validates that structured prompting techniques map directly to multi-agent debate architectures, enabling equivalent outcomes through structural equivalence.

Does cognitive diversity alone improve multi-agent ideation quality?

Multi-agent teams substantially outperform solo ideation, but only when members possess genuine senior knowledge. Diverse teams without expertise underperform even a single competent agent, because cognitive stimulation without expertise triggers process losses instead of insight.

Can abstractions guide exploration better than depth alone?

RLAD jointly trains abstraction and solution generators, showing that allocating test-time compute to diverse abstractions outperforms parallel solution sampling at large budgets. Abstractions create structured breadth-first exploration that prevents the underthinking failure mode of depth-only reasoning chains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.