INQUIRING LINE

Why can brilliant researchers look at the same AI output and disagree on whether it's truly original?

Why do top scientists disagree on whether o3 produces genuinely novel ideas?

This explores why expert judgment splits on whether a frontier reasoning model like o3 comes up with ideas that are actually new. The library has no notes about o3 specifically, but it explains well why this kind of disagreement keeps happening with any model.


This explores why expert judgment splits on whether a frontier reasoning model like o3 comes up with ideas that are actually new. To be clear up front, the collection has no notes about o3 itself or about particular scientists' public statements on it. What it does have is a clear account of why reasonable experts can look at the same AI output and disagree. The short version: the disagreement may not come from anyone being wrong. It comes from the fact that 'novel' isn't something you can measure the way you measure accuracy. Parker argues that claims of AI discovery will stay contested for years, because whether an idea counts as novel and useful is decided by a research community, not by an objective test. Mathematics may be the one exception, because a computer can check a proof Will we ever agree on whether AI makes real discoveries?.

The experimental evidence shows that both sides can be right at once. In a large study of more than 100 NLP researchers, experts rated LLM-generated research ideas as significantly more novel than ideas written by human experts Do language models generate more novel research ideas than experts?. The explanation is surprising. Experts are held back by what they know won't work, while models freely combine concepts across fields without that restraint Can LLMs generate more novel ideas than human experts?. But when 43 researchers spent over 100 hours each actually carrying out randomly assigned ideas, the AI ideas lost far more quality than the human ones. Evaluation plans turned out to be impractical, and technical groundwork was missing Do LLM research ideas actually hold up when experts try to execute them?. So a scientist who judges ideas on paper may see real novelty, while one who judges by what survives in the lab may see mostly noise.

There is a second split: the models generate well but judge poorly. They produce surprising combinations but avoid committing to whether an idea is feasible or valid. Automated evaluation overestimates idea quality by about 60% Why do LLMs generate more novel research ideas than experts?. That means much of the 'is it real?' work falls on human experts, and each brings their own standards. Even favorable evidence often comes from the systems' own builders. Co-Scientist's team reports that hypotheses get better as the system is given more compute, but replication so far comes mostly from their own validation Does more thinking time improve AI-generated research hypotheses?. Meanwhile, MIT's withdrawal request for an influential AI-and-science preprint shows how a claim can shape the debate long before anyone has checked it Can unreviewed preprints shape scientific debate before peer review?.

Two deeper points help explain why the debate is so hard to settle. First, matching outputs don't prove matching understanding. Networks can produce identical results while their internal representations are fractured in ways that block creative recombination in new settings Can identical outputs hide broken internal representations?. A striking idea might come from real generalization, or from lucky surface-level recombination, and the text alone can't show which. Second, in science, much of an idea's weight comes from who proposes it: their track record and their stake in the field. Models work without that social context Can language models distinguish expert arguments from common assumptions?. Part of what top scientists disagree about may be whether an idea can count as a discovery when no one stands behind it.


Sources 9 notes

Will we ever agree on whether AI makes real discoveries?

Parker argues the debate over AI-generated discoveries will persist for years because 'novel' and 'useful' depend on community judgment rather than objective criteria. Mathematics may be the sole exception, since theorems can be rigorously verified by computer.

Do language models generate more novel research ideas than experts?

A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.

Can LLMs generate more novel ideas than human experts?

LLMs produce more novel research ideas than experts because they lack disciplinary constraints, but they systematically avoid evaluative stance-taking required to assess feasibility or validity. Generation and evaluation are dissociated capabilities.

Do LLM research ideas actually hold up when experts try to execute them?

When 43 expert researchers implemented randomly-assigned ideas over 100+ hours, LLM-generated ideas declined significantly more than human ideas across all metrics. Execution revealed systematic weaknesses invisible at ideation, including impractical evaluation designs and missing technical groundwork.

Why do LLMs generate more novel research ideas than experts?

Research shows LLM-generated ideas are statistically more novel than expert-produced ideas, but LLMs struggle to evaluate quality—automated evaluation overestimates by 60%. When executed, LLM ideas drop significantly on all metrics, suggesting novelty without feasibility.

Show all 9 sources
Does more thinking time improve AI-generated research hypotheses?

Co-Scientist's builders report that Elo ratings of AI-generated hypotheses improve as the system spends more compute in a tournament-based debate loop. Three biomedical cases show independent recapitulation and testable predictions, though replication is limited to the builders' own validation.

Can unreviewed preprints shape scientific debate before peer review?

MIT's case demonstrates that an arXiv preprint shaped AI and science discussions extensively despite never undergoing peer review. When the institution later raised reliability concerns, the damage to discourse had already occurred.

Can identical outputs hide broken internal representations?

Networks trained with SGD reproduce outputs perfectly while having radically different internal structure than evolved networks, with weight perturbations revealing fractured, entangled representations that prevent transfer to novel contexts or creative recombination.

Can language models distinguish expert arguments from common assumptions?

LLMs lose the social context that gives expert claims their force—reputation, track record, and standing—because they process only text, not the social world where expertise is built and evaluated.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.