SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

Can evolutionary search discover machine learning algorithms from scratch?

Can algorithms for learning and prediction be discovered through evolutionary search over basic mathematical operations, without human-designed components like neural networks or backpropagation?

Synthesis note · 2026-10-08 · sourced from Frontier AI Risk & RSI

AutoML-Zero sets out to "automatically search for whole ML algorithms using little restriction on form and only simple mathematical operations as building blocks," representing each algorithm as a program with three component functions — Setup, Predict, and Learn — built from a vocabulary of 65 basic ops ("those that are typically learned by high-school level," excluding "machine learning concepts, matrix decompositions, and derivatives"). Starting from "empty programs," evolutionary search discovers algorithms that are first competitive with "two-layer neural networks trained by backpropagation," then, evolved directly on CIFAR-10 variants, produce "modern techniques" including bilinear interactions, normalized gradients, and weight averaging — measured against hand-designed baselines on held-out CIFAR-10 test data and shown to generalize to SVHN, ImageNet, and Fashion MNIST.

The paper is explicit that genericity trades against searchability: "the space is so generic that it ends up being quite sparse," with good algorithms for even a trivial task "as rare as 1 in 10^12," so random search fails and the authors build infrastructure capable of searching "10,000 models/second/cpu core," plus a functional-equivalence-checking technique that skips re-evaluating algorithms "that have already been seen, even if they have different implementations," for a "4x speedup." Selection runs by tournament — each cycle samples a subset of the population, keeps the best as a parent, mutates it into a child, discards the oldest member — with no gradient signal and no specified loss landscape, only accuracy on proxy tasks. Three controlled experiments then show the evolved algorithms adapting to task conditions rather than converging on one fixed solution: with only 80 training examples over 100 epochs, a noisy-ReLU-like regularizer reminiscent of Dropout emerges reproducibly (8/30 repeats vs. 0/30 in an 800-example control, p<0.0005); with 10 epochs instead of 100, learning-rate decay emerges almost universally (30/30 vs. 3/30 control, p<10⁻¹⁴); and across all 10 CIFAR-10 classes instead of a binary split, algorithms preferentially use the transformed weight-matrix mean as a learning rate (24/30 vs. 0/30, p<10⁻¹¹).

AutoML-Zero sits at the earlier end of the lineage that Can automated scoring verify mathematical constructions without human understanding? extends to mathematical construction search: both pair evolved programs with a fully automated score — here, accuracy on proxy and held-out tasks; there, a verifier's score — rather than a proof of correctness, and neither establishes that the evolved program itself is interpretable. AutoML-Zero's own limitations section names the same gap from the inside: tuning evolved constants is "insufficient due to hyperparameter coupling," and without inspection "we may not know what each variable means." Against Can autonomous research pipelines discover AI architectures that AutoML cannot?, the scope here is narrower by design: AutoML-Zero's target is exactly the hyperparameter/architecture-search family that the later work contrasts with autoresearch's code comprehension and bug diagnosis. AutoML-Zero reduces human bias inside that family — no hand-designed layers or backprop rule assumed — but it still searches a fixed, if generic, instruction space; it cannot read a data pipeline or diagnose a bug the way the later autoresearch system does. Can language models discover new expertise through collaborative weight search? occupies an adjacent point in the same design space — training-free, assumption-free population search — but over continuous weight space rather than discrete program instructions, and without AutoML-Zero's build-from-empty-program ambition.

The excerpt does not give an interpretability analysis of the final evolved algorithms beyond the handful discussed in Discussion, does not report compute cost against its own 1-in-10^12 sparsity estimate, and does not address whether the approach scales past CIFAR-10-sized proxy tasks to larger modern architectures — the authors flag this directly ("there is still much work to be done"). The implication, at the strength the evidence allows, is that removing hand-designed building blocks from a search space does not by itself solve search: it was the infrastructure — parallelism, functional equivalence checking, diversity, hurdles — that made the generic space tractable, suggesting that reducing human bias in a search space and making that space searchable are two separate engineering problems, not one.

Inquiring lines that read this note 5

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI systems discover fundamental improvements to their own architectures? Can AI agents improve their skills through accumulated experience and reuse? Does AI-assisted research sacrifice exploration breadth for productivity gains? Can AI systems achieve real improvement without external human feedback?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 109 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

AutoML-Zero evolves neural networks, gradient descent and weight averaging from basic math operations with no human-designed building blocks