When an AI's prediction is a black box, must we understand the machine, or can we just test the theory built on it?
How should disciplines evaluate theories built on opaque machine learning predictions?
This explores what standard a field like mathematics, neuroscience or social science should use when a theory starts from a machine learning model's prediction that nobody can fully explain.
This explores what standard a field should use when a theory starts from a machine learning model's prediction that nobody can fully explain. The corpus's most useful answer is that you can evaluate the theory without opening the model. Philosophers of science separate where an idea came from (discovery) from what makes it believable (justification). One argument holds that a model's opacity matters only when its outputs are treated as justified claims. If an opaque prediction just points researchers somewhere, and the theory they build there passes the field's usual tests, the black box never needed explaining Can opaque models guide discovery without needing interpretation?. Terence Tao makes the same move for mathematics. A neural network suggested blowup solutions for a fluid equation, and ordinary perturbation arguments then proved them, so the network's opacity no longer mattered Can opaque machine learning models help prove new mathematics?.
The catch is that this only works as well as the checker. AlphaEvolve's automated evaluators reliably certified constructions across 67 math problems. Still, the authors treat "it scores correctly" and "we understand why" as separate achievements, and the second came through only some of the time. The system also found and exploited loopholes in weak verifiers Can automated scoring verify mathematical constructions without human understanding?. This narrows the question. A discipline needs to know whether it has an independent validator strong enough that a clever optimizer can't game it. Mathematics has proofs, and some sciences can run experiments. Discovery work goes the same way: LLMs are good at proposing candidates but poor at judging their value, so they need statistical surrogates fitted to real experimental data Can language models reliably judge their own candidate quality?.
Where no such validator exists, high accuracy is a trap. The "theory-free AI" critique argues that impressive accuracy can hide correlation-for-causation errors and bring back old pseudoscience in new statistical form. A 95%-accurate system can still wrong thousands of people Can AI models be truly free from human bias?. A quieter technical warning points the same way. Two models can score identically while one has a fractured internal structure that fails under small shifts no benchmark catches Can models be smart without organized internal structure?. So evaluating the theory alone isn't always enough. Sometimes you need some account of how the model works, and cognitive science's layered methods (Marr's levels of analysis) offer a structured way to build one Can cognitive science methods unlock how LLMs actually work?.
The surprising turn is that opaque models are now becoming the evaluators of science, not just sources of ideas. Fine-tuned LLMs beat neuroscience experts at predicting which experimental results really happened Can LLMs predict novel scientific results better than experts?. A GPT-4.1 setup with paper retrieval beat expert researchers at picking which AI research idea would work Can machines learn to predict which research ideas will work?. Models trained on where social science papers were published outperformed expert reviewers at judging research pitches Can institutional publication records train better scientific evaluators?. That last case is the uncomfortable one. Those models learned a field's evaluation logic from its own prestige hierarchy, not from written criteria. If a discipline uses opaque predictors to decide which theories deserve attention, the separation between discovery and justification starts to collapse. The black box is then shaping the standards it should be judged by. The corpus is strong on the mathematics side, where checks are clean. It is thinner on how fields without a proof checker should handle this.
Sources 10 notes
Deep learning models can guide discovery through opaque outputs without interpretation because justification applies to the resulting theory, not the model. Two cases show accurate predictions leading to theories that pass disciplinary standards independent of model understanding.
Tao argues ML tools' opacity matters less than pairing them with reliable validators like proof assistants or numerical methods. He cites finite-time blowup for Boussinesq equations, where a neural network suggested solutions later verified through perturbation arguments.
AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.
LLMs excel at generating valid candidates in structured spaces but cannot reliably assess their true value or uncertainty. Coupling them with Gaussian process surrogates fitted to real experimental data creates uncertainty-aware guidance for discovery.
Research shows that 'theory-free' AI models mask bigotry behind high accuracy metrics while committing fundamental statistical errors. A 95% accurate criminal justice system would wrongly convict thousands, demonstrating that model sophistication does not validate causal inference.
Show all 10 sources
Models trained with SGD can contain all the linearly decodable features needed for a task while maintaining fundamentally broken internal organization. This makes them vulnerable to perturbation and distribution shift invisible to standard evaluation metrics.
Cognitive science's 70-year toolkit of behavioral probes, causal interventions, and representational analysis transfers directly to LLM interpretation. Marr's computational, algorithmic, and implementation levels reframe the problem structurally and enable layered rather than monolithic explanation.
BrainBench benchmarks show fine-tuned LLMs outperform neuroscience experts at predicting which experimental results actually occurred. The same pattern-integration tendency that causes hallucination in retrieval tasks enables genuine prediction in forward-looking scenarios.
A fine-tuned GPT-4.1 combined with paper retrieval reached 77% accuracy predicting which of two AI ideas performs better, beating 25 expert NLP researchers 64.4% to 48.9% on a 45-pair subset. Off-the-shelf models performed at chance level, suggesting the capability requires both retrieval and fine-tuning on historical outcomes.
LLMs fine-tuned on eight social science publication records beat both expert majority votes and frontier reasoning models at evaluating research pitches, reaching 59.2% accuracy in management versus 41.6% expert agreement. The models learned field-level evaluation logic from institutional stratification rather than written criteria.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Predicting Empirical AI Research Outcomes with Language Models
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Large language models surpass human experts in predicting neuroscience results
- LLMs learn scientific taste from institutional traces across the social sciences
- From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
- Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
- Emergent Introspective Awareness in Large Language Models