Line of inquiry
Inquiring lines›Why are language models fragile de…›Why are LLM outputs so inconsisten…›this line of inquiry
Why do LLMs generate novel ideas but struggle to evaluate them?
A broader line of inquiry — a family of 37 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 37
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What makes novelty assessment harder to automate than idea generation?
- Can LLMs generate more novel research ideas than human experts?
- Why do research ideation systems suffer from diversity collapse despite high novelty metrics?
- Why do LLMs generate novel ideas but struggle to evaluate them?
- Why do LLM-generated ideas score higher novelty yet lower feasibility than expert ideas?
- Can LLMs reliably assess the quality of ideas they generate?
- Why do LLM research ideas lack diversity despite high average novelty?
- How can LLMs evaluate their own creative outputs for utility and novelty?
- Do LLMs generate more novel ideas than they can evaluate?
- Can human researchers improve LLM ideas through iterative feedback?
- Why does LLM research ideation collapse into low diversity despite high novelty?
- Why does diversity collapse occur in multi-agent research ideation despite high novelty?
- Why do LLMs generate ideas that sound novel but fail during execution?
- Can LLM diversity collapse in research ideation be reversed or mitigated?
- How do constrained versus unconstrained domains flip LLM novelty patterns?
- Why do LLMs generate novel ideas but lack evaluative commitment?
- Do novelty and feasibility always trade off in idea generation?
- Why do models generate creative ideas but fail to evaluate their legitimacy?
- Why are AI research ideas more novel but harder to evaluate than human ones?
- Can structured decomposition fix evaluation gaps in other research tasks?
- Can few-shot examples narrow generative diversity in creative tasks?
- How should AI ideation systems decompose and recombine research concepts?
- Why do LLMs plateau on creativity tasks while humans reach further?
- What graph structures would enable transformational creative reasoning in LLMs?
- What distinguishes scientific plausibility from cognitive availability in research ideas?
- Can structured evaluation assess novelty in scientific writing?
- Can proxy evaluation of ideas accurately predict their quality without implementation?
- Which LLM backends produce the most executable research ideas?
- What makes a novel research idea practically infeasible for implementation?
- Why does automated evaluation consistently overestimate research quality?
- Why do some LLM clusters cite broader psychology than others?
- How do expert priors constrain human researchers from exploring novel concepts?
- Can ranking by coherence while minimizing author-community coverage find novel research?
- Do LLMs match top human creative writers in literary quality?
- How does prompt design alter what kind of creativity LLMs can express?
- How does the Word Novelty Rate metric measure convention formation?
- Why do excerpts and abstracts miss the recovery ideas present in full papers?