Will we ever agree on whether AI makes real discoveries?
Can AI systems produce genuine scientific breakthroughs, or will the field remain divided over what counts as discovery? The answer may depend on whether subjective human judgment can be replaced by objective verification.
Edward Parker, writing for AI Frontiers, argues that whether large language models can produce genuine new scientific discoveries will stay unresolved "for years," because the field lacks a workable test for what counts as a "scientific discovery by an AI." He contrasts earlier reinforcement-learning successes like AlphaFold — "essentially computational in nature," solving "well-understood computational math problems" — with the much higher bar of "a new conceptual scientific idea," the kind that prompts a person to say "I get it now" and build further insight on it, as with Darwin's natural selection or Einstein's spacetime curvature. He notes OpenAI president Greg Brockman's claim that o3 and o4-mini are "the first models where top scientists tell us they produce legitimately good and useful novel ideas," but says he has not personally found an LLM example of synthesizing training data into a new conceptual discovery comparable to Kekulé's benzene-ring insight.
Parker's reasoning is that "novel" and "useful" are not objective properties a model's output can be checked against, but community judgments: "what the research community considers interesting and important is more arbitrary than they'd like to admit," resting partly on "subjective notions of beauty or whether the topic happens to be in vogue." Producing a correct new result is easy — "combining several well-known facts in a simple logical chain" — but recognizing which correct results matter is a judgment call, so any AI claim to discovery will split experts into those calling it groundbreaking and those calling it obvious, with "nonexperts" unable to adjudicate. He singles out pure mathematics as the exception, since a new theorem "can, with effort, be expressed precisely enough to be rigorously verified entirely by computer," giving math an objective check that experimental science and conceptual insight lack.
This sits against the framing in Can artificial intelligence ever truly understand science?: Parker's prediction of a lasting "messy middle ground" is the same observation made from the outside — no one has yet witnessed the "agent" case clearly enough to end the argument. It also bears on the concrete test case in Can AI systems generate hypotheses that match unpublished experimental discoveries?: Parker's framework predicts that even a hit like this will not settle the debate, since "how much credit (if any) belongs to the LLM" is exactly the kind of subjective attribution he says resists adjudication. And it complements the structural case in What stops AI from discovering science without human help?, which argues from inside AI research that the gaps are architectural; Parker reaches the same stalemate from the sociology of scientific judgment rather than from model design.
The excerpt is an opinion essay, not a study — Parker offers no survey, benchmark, or dataset, only his own reading of cases like AlphaFold and Brockman's claim, and he explicitly says he hopes to be proven wrong. It does not establish that current LLMs cannot make conceptual discoveries, only that the criteria for recognizing one are unsettled; the piece's own suggestion — that formally verifiable math theorems are the "lowest-hanging fruit" for an unambiguous AI discovery — points toward where the first clear test case is most likely to appear, sooner than in experimental or conceptual science.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do philosophical assumptions about AI consciousness affect practical harms and design? Why do LLM research ideation systems generate novelty but lack diversity? What human oversight must AI research systems have?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can artificial intelligence ever truly understand science?
Researchers ask whether AI can move beyond predicting outcomes to genuinely grasping the theories behind them. The question hinges on what scientific understanding actually means.
Parker's predicted "messy middle ground" echoes this paper's finding that no true agent case has yet been observed
-
Can AI systems generate hypotheses that match unpublished experimental discoveries?
Researchers tested whether an AI hypothesis-generation platform could arrive at mechanisms their own labs had experimentally confirmed but not yet published, exploring whether AI can independently discover known-but-hidden biological answers.
gives the kind of concrete case Parker says will fuel endless disagreement over AI credit
-
What stops AI from discovering science without human help?
Can current agentic AI systems autonomously conduct natural-science discovery, or do fundamental gaps in training and deployment block them? This matters because it shapes realistic expectations for AI in research.
reaches the same stalemate from AI research design rather than from the sociology of scientific judgment
-
Can opaque models guide discovery without needing interpretation?
Does deep learning need to be interpretable when it steers hypothesis formation rather than standing as a justified claim itself? The distinction matters for when opacity becomes an epistemic problem.
Extends: the discovery/justification distinction explains why math's verifiability, unlike other fields, makes AI-assisted discovery claims uncontested
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- We'll Be Arguing for Years Whether Large Language Models Can Make New Scientific Discoveries
- Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- LLMs learn scientific taste from institutional traces across the social sciences
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- Deep Learning Opacity in Scientific Discovery
- Mathematical methods and human thought in the age of AI
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Original note title
Parker argues scientific discovery by AI will remain contested because novel and useful are subjective judgments — mathematics may be the exception