INQUIRING LINE

When AI evolves its own algorithms instead of a human writing them, who actually understands how they work?

What interpretability challenges arise when algorithms are discovered rather than designed?

This explores what makes algorithms hard to understand when a search process finds them (evolutionary search, AI-driven discovery, or gradient descent itself) instead of a person writing them step by step, and whether we always need to understand them.


This explores what makes algorithms hard to understand when a search process finds them instead of a human writing them, and whether we always need to understand them. The corpus suggests a surprising answer: with discovered algorithms, *checking that something works* and *understanding why it works* come apart, and much of the interesting research lives in the gap between the two.

Start with how strange discovered algorithms can be. AutoML-Zero began with nothing but 65 basic math operations and used evolutionary search to rebuild neural networks, gradient descent, and tricks like weight averaging and learning-rate decay Can evolutionary search discover machine learning algorithms from scratch?. It also adjusted its strategies to suit each task. Nobody designed these programs, so nobody starts out knowing why a given line is there. AlphaEvolve pushes this further. Automated evaluators kept an evolutionary loop going long enough to produce faster algorithms and better hardware designs Can machine feedback sustain discovery at test time?. The authors make a careful distinction: the evaluator's score reliably certified solutions across 67 math problems, but humans or tools could explain those solutions only some of the time Can automated scoring verify mathematical constructions without human understanding?. There was a sharper problem too. The search sometimes exploited loopholes in the verifier. When you can't read the algorithm, you can't easily tell whether it solved the problem or gamed the test.

The less obvious point is that "it passes the test" can hide a mess underneath. One line of research shows that networks with identical accuracy can have very different internal organization. Some are cleanly structured and others are fractured, and the fractured ones fail under perturbation or distribution shift in ways standard metrics never reveal Can models be smart without organized internal structure?. Any neural network trained by gradient descent is itself a discovered algorithm, so the AlphaEvolve problem is really the everyday problem of deep learning, just made more visible.

So what do people do about it? One camp argues that opacity matters less than it seems. Terence Tao's position is that an opaque model is fine as long as its output is passed to a reliable validator, such as a proof assistant or a rigorous numerical check. He points to a case where a neural network suggested blowup solutions to a fluid equation that mathematicians later proved by hand Can opaque machine learning models help prove new mathematics?. A philosophy-of-science version of this argument holds that opacity only becomes a problem when you treat the model's output as the justification itself. If the model just points you somewhere and the resulting theory then passes ordinary standards, the black box never needed opening Can opaque models guide discovery without needing interpretation?. The other camp tries to build understanding back in. Training with sparse weights forces networks into small circuits where neurons map to simple concepts, though this hasn't yet scaled past tens of millions of parameters Can sparse weight training make neural networks interpretable by design?. A third option is to have an LLM explain the discovered artifact in plain language. This is powerful, but it brings its own risk: explanations that sound right but don't faithfully describe what the algorithm does Can natural language explanations redefine what interpretability means?.

The takeaway is that for discovered algorithms, the verifier becomes the thing you most need to trust. Whether you can live without understanding depends on how strong your checker is and whether you'll use the result as a lead to follow or as proof that something is true. When the checker is weak, the search process will find its blind spots before you do.


Sources 8 notes

Can evolutionary search discover machine learning algorithms from scratch?

AutoML-Zero evolved algorithms from 65 basic operations that match neural networks and rediscover modern techniques like weight averaging and learning-rate decay, adapting strategies to task conditions in controlled experiments.

Can machine feedback sustain discovery at test time?

AlphaEvolve demonstrates that automated evaluators can sustain evolutionary loops long enough to produce real discoveries—faster algorithms, optimized hardware designs, and improved training methods. The key is that cheap, objective verification closes the generation-verification gap where discovery becomes computationally feasible.

Can automated scoring verify mathematical constructions without human understanding?

AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.

Can models be smart without organized internal structure?

Models trained with SGD can contain all the linearly decodable features needed for a task while maintaining fundamentally broken internal organization. This makes them vulnerable to perturbation and distribution shift invisible to standard evaluation metrics.

Can opaque machine learning models help prove new mathematics?

Tao argues ML tools' opacity matters less than pairing them with reliable validators like proof assistants or numerical methods. He cites finite-time blowup for Boussinesq equations, where a neural network suggested solutions later verified through perturbation arguments.

Show all 8 sources
Can opaque models guide discovery without needing interpretation?

Deep learning models can guide discovery through opaque outputs without interpretation because justification applies to the resulting theory, not the model. Two cases show accurate predictions leading to theories that pass disciplinary standards independent of model understanding.

Can sparse weight training make neural networks interpretable by design?

Training transformers with sparse weights creates compact, human-interpretable circuits where neurons correspond to simple concepts with clear connections. Ablation studies confirm these circuits are necessary and sufficient for task performance, though scaling beyond tens of millions of parameters while maintaining interpretability remains unsolved.

Can natural language explanations redefine what interpretability means?

LLMs' capacity to explain in natural language expands the scale and complexity of patterns conveyable to humans, enabling ambitious new interpretability goals including model-to-model auditing. However, this medium introduces critical risks: hallucinated explanations that feel plausible but lack faithfulness.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.