SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

Will we ever agree on whether AI makes real discoveries?

Can AI systems produce genuine scientific breakthroughs, or will the field remain divided over what counts as discovery? The answer may depend on whether subjective human judgment can be replaced by objective verification.

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

Edward Parker, writing for AI Frontiers, argues that whether large language models can produce genuine new scientific discoveries will stay unresolved "for years," because the field lacks a workable test for what counts as a "scientific discovery by an AI." He contrasts earlier reinforcement-learning successes like AlphaFold — "essentially computational in nature," solving "well-understood computational math problems" — with the much higher bar of "a new conceptual scientific idea," the kind that prompts a person to say "I get it now" and build further insight on it, as with Darwin's natural selection or Einstein's spacetime curvature. He notes OpenAI president Greg Brockman's claim that o3 and o4-mini are "the first models where top scientists tell us they produce legitimately good and useful novel ideas," but says he has not personally found an LLM example of synthesizing training data into a new conceptual discovery comparable to Kekulé's benzene-ring insight.

Parker's reasoning is that "novel" and "useful" are not objective properties a model's output can be checked against, but community judgments: "what the research community considers interesting and important is more arbitrary than they'd like to admit," resting partly on "subjective notions of beauty or whether the topic happens to be in vogue." Producing a correct new result is easy — "combining several well-known facts in a simple logical chain" — but recognizing which correct results matter is a judgment call, so any AI claim to discovery will split experts into those calling it groundbreaking and those calling it obvious, with "nonexperts" unable to adjudicate. He singles out pure mathematics as the exception, since a new theorem "can, with effort, be expressed precisely enough to be rigorously verified entirely by computer," giving math an objective check that experimental science and conceptual insight lack.

This sits against the framing in Can artificial intelligence ever truly understand science?: Parker's prediction of a lasting "messy middle ground" is the same observation made from the outside — no one has yet witnessed the "agent" case clearly enough to end the argument. It also bears on the concrete test case in Can AI systems generate hypotheses that match unpublished experimental discoveries?: Parker's framework predicts that even a hit like this will not settle the debate, since "how much credit (if any) belongs to the LLM" is exactly the kind of subjective attribution he says resists adjudication. And it complements the structural case in What stops AI from discovering science without human help?, which argues from inside AI research that the gaps are architectural; Parker reaches the same stalemate from the sociology of scientific judgment rather than from model design.

The excerpt is an opinion essay, not a study — Parker offers no survey, benchmark, or dataset, only his own reading of cases like AlphaFold and Brockman's claim, and he explicitly says he hopes to be proven wrong. It does not establish that current LLMs cannot make conceptual discoveries, only that the criteria for recognizing one are unsettled; the piece's own suggestion — that formally verifiable math theorems are the "lowest-hanging fruit" for an unambiguous AI discovery — points toward where the first clear test case is most likely to appear, sooner than in experimental or conceptual science.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do philosophical assumptions about AI consciousness affect practical harms and design? Why do LLM research ideation systems generate novelty but lack diversity? What human oversight must AI research systems have?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 95 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Parker argues scientific discovery by AI will remain contested because novel and useful are subjective judgments — mathematics may be the exception