INQUIRING LINE

Why does a computer searching from scratch, with no human hints about what works, mostly find useless junk?

Why do good algorithms become rarer as the search space grows more generic?

This explores why, when you let a computer search for algorithms with very few built-in human assumptions (a 'generic' search space built from basic math operations rather than pre-made neural network parts), working solutions become needles in an exponentially larger haystack, and what the corpus shows researchers doing about it.


This explores why a search space with fewer human assumptions built in makes good algorithms harder to find, and how systems in the corpus cope with that. The short version is arithmetic. If your building blocks are already high-level parts like layers, optimizers and loss functions, almost any combination does something sensible. If your building blocks are raw operations like add, multiply and read from memory, the number of possible programs grows exponentially with length, while the number that actually learn anything barely grows. Most random programs in a generic space are noise. That is the price of not telling the search what a good answer looks like. The clearest case in the collection is Can evolutionary search discover machine learning algorithms from scratch?. Starting from 65 basic math operations, evolutionary search still rediscovered neural networks, gradient descent, weight averaging and learning-rate decay. It got there by keeping and mutating the rare programs that showed any signal, not by sampling blindly. The corpus doesn't measure exactly how sparse good solutions get as the space generalizes. What it does show is the methods people build so they don't have to rely on luck.

The surprising lesson is that adding more randomness doesn't help. You might expect wider, noisier exploration to find more needles. But Does adding randomness alone improve recursive reasoning models? found that injecting raw randomness into reasoning models produced no gains. Improvement came only when the randomness was tied to a principled objective. The same holds in algorithm search: in a sparse space, undirected noise just samples more of the haystack. The signal that guides the search matters more than how much ground it covers. That is why Can tree search replace human feedback in LLM training? is useful: tree search ranks partial paths by whether they lead to success, which turns a flat search into a guided one.

A second answer is to stop treating the search method as fixed. Can an AI system improve its own search methods automatically? puts an outer loop on top of the search. It reads the inner search code, spots where it gets stuck in repetitive patterns, and writes new search mechanisms at runtime, such as bandit methods and combinatorial optimization. The result was a 5x improvement. Can AI systems improve themselves through trial and error? takes a related route. It keeps an archive of agent variants instead of only the current best one, so mediocre intermediate versions can serve as stepping stones toward better ones. In sparse spaces the path to a good solution often runs through things that don't look good yet.

The opposite failure is just as instructive. If the search narrows too fast, it stops looking for rare solutions altogether. Does reinforcement learning squeeze exploration diversity in search agents? shows that reinforcement learning collapses search agents onto a few reward-maximizing habits. Does preference tuning always reduce diversity the same way? finds the same convergence in code, where training pushes toward one 'correct' style. Inside a single model's chain of thought, Why do reasoning models abandon promising solution paths? describes both problems together: aimless wandering, and dropping promising paths too early. Searching a large space well means balancing breadth and commitment, and the needle is easy to miss in either direction.

The scoring rule is the last piece. When good solutions are rare, the search drifts toward whatever the evaluator rewards, and that may not be what you wanted. Why do fixed benchmarks fail as agents grow stronger? argues that fixed benchmarks get gamed as search gets stronger, and proposes changing the objective at set intervals. So the real tradeoff in generic search spaces is this: removing human assumptions opens the door to genuinely new algorithms, but it moves all the burden onto how you guide, diversify and score the search. AutoML-Zero rediscovering gradient descent from scratch shows the door is real. It also shows that most of the engineering effort goes into the search, not the space.


Sources 9 notes

Can evolutionary search discover machine learning algorithms from scratch?

AutoML-Zero evolved algorithms from 65 basic operations that match neural networks and rediscover modern techniques like weight averaging and learning-rate decay, adapting strategies to task conditions in controlled experiments.

Does adding randomness alone improve recursive reasoning models?

GRAM's ablations show naive stochasticity added to existing models yields no improvement. Gains come specifically from amortized variational inference, which couples stochastic latents to a principled generative objective rather than injecting undirected noise.

Can tree search replace human feedback in LLM training?

AlphaLLM uses tree search outcomes and three critic models to derive dense reward signals equivalent to human-labeled feedback. Tree structure naturally ranks solution paths by success, replacing the annotation oracle that standard RLHF requires.

Can an AI system improve its own search methods automatically?

An outer loop successfully read inner loop code, identified bottlenecks, and generated new Python mechanisms at runtime, discovering combinatorial optimization and bandit methods that broke the inner loop's deterministic patterns and improved performance on GPT pretraining by 5x.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

Show all 9 sources
Does reinforcement learning squeeze exploration diversity in search agents?

RL training compresses behavioral diversity in search agents through the same entropy collapse mechanism documented in reasoning—policies converge on narrow reward-maximizing strategies. SFT on diverse demonstrations preserves exploration breadth, suggesting diversity-preservation techniques are essential for RL search scaling.

Does preference tuning always reduce diversity the same way?

RLHF reduces lexical-syntactic diversity in code generation but increases it in creative writing. The direction depends on what each domain incentivizes: code rewards convergence toward correct solutions, while creative writing rewards stylistic distinctiveness.

Why do reasoning models abandon promising solution paths?

Reasoning LLMs exhibit two reinforcing failures: wandering (invalid exploration) and underthinking (premature path-switching). Decoding-level interventions like thought-switching penalties improve accuracy without fine-tuning, suggesting viable solutions exist but are abandoned prematurely.

Why do fixed benchmarks fail as agents grow stronger?

Static benchmarks saturate and invite gaming as agents strengthen. RQGM solves this by splitting search into epochs with fixed criteria per epoch but evolving objectives across boundaries, keeping improvement guarantees while moving the target faster than agents can exploit it.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.