Can an AI that breeds, mutates and scores its own candidate answers while it thinks go beyond planning puzzles?
Can evolutionary search at inference time scale beyond natural language planning?
This explores whether the evolutionary approach, where an LLM breeds, mutates and selects candidate answers while it runs, only works for natural-language planning puzzles or can carry over to other kinds of problems. The corpus has one direct result plus several nearby findings that point to where it might stretch and where it might hit a wall.
This explores whether evolutionary search at inference time (an LLM generating a population of candidate solutions, recombining and mutating the promising ones, and repeating) is a trick for planning tasks or a more general way to spend compute. The direct evidence comes from Mind Evolution, which solves over 98% of benchmark planning tasks and clearly beats both best-of-N sampling and step-by-step self-revision Can evolutionary search beat sampling and revision at inference time?. A key detail is that it works directly in natural language. Nobody has to translate the problem into a formal solver language first. That is what makes it look portable. What it does still need is some way to score each candidate, and that requirement turns out to be the real limit on how far it can scale.
The most interesting sign that it can go further comes from a different direction: agents that evolve the scoring itself. SAGA runs two loops. An outer LLM loop proposes new objectives in plain language and compiles them into executable scoring functions, and an inner loop optimizes against them Can agents evolve their own objectives during search?. In practice, "what counts as a good answer" becomes something the search can discover instead of a fixed input. That points to how evolutionary search might move into open-ended areas like scientific discovery: when there is no ready-made checker, the system writes and revises its own.
There is also a less obvious reason the outer loop matters. LLMs are poor at running iterative procedures in their heads. When asked to carry out a numerical optimization, they pattern-match a remembered solution and output plausible but wrong numbers Do large language models actually perform iterative optimization?. On real constrained-optimization problems they level off at about 55–60% of constraints satisfied, whatever their size or training Do larger language models solve constrained optimization better?. Read alongside Mind Evolution, this suggests a division of labor. The model is good at proposing and recombining ideas but bad at iterating, so an external evolutionary loop supplies the iteration the model can't do internally. If that reading holds, the domains where evolutionary search should help most are these hard optimization problems, as long as candidates can be checked.
Evolutionary search also fits a broader pattern: inference-time compute keeps showing new axes you can scale. Agentic research systems show that the number of search iterations scales much like reasoning tokens, with steady gains that eventually flatten Does search budget scale like reasoning tokens for answer quality?. Population-based search is plausibly another axis of this kind. The caution is that search only amplifies what the model can already propose. Reasoning models stay ahead of non-reasoning ones however much inference compute the latter get Can non-reasoning models catch up with more compute?, and reasoning tends to break on unfamiliar instances rather than long ones Do language models fail at reasoning due to complexity or novelty?. If every candidate in the population comes from the same blind spot, recombining them won't get you out of it.
The honest gap is that the corpus has no paper testing Mind Evolution itself outside planning. The case for scaling beyond planning is assembled from neighboring findings, not shown directly. The takeaway is this: evolutionary search seems to scale as far as your ability to score candidates, and the most promising frontier is systems that evolve the scorer along with the solutions.
Sources 7 notes
Mind Evolution, an evolutionary search strategy using LLM-generated crossover and mutation with island model diversity, solves 98%+ of planning tasks and significantly outperforms best-of-N and sequential revision strategies while working directly in natural language without task formalization.
SAGA's bi-level architecture closes a feedback loop from optimization results back to goal design by having an outer LLM loop propose new objectives and compile them into code the inner loop can immediately use, enabling objective formulation as part of discovery rather than a fixed input.
Research shows LLMs cannot perform iterative procedures in latent space. They recognize optimization problems as template-similar and emit plausible-looking but incorrect values, a failure mode that persists across model scale and training approaches.
Across constrained-optimization tasks, LLMs converge to ~55–60% constraint satisfaction independent of architecture, parameter count, or training regime. Reasoning models do not systematically outperform standard models, suggesting a fundamental ceiling rather than a scaling gap.
Agentic deep research shows monotonic-to-diminishing-returns curves for search iterations, matching reasoning token scaling. This creates a new inference-compute axis: models can trade off reasoning budget against search budget to optimize answer quality.
Show all 7 sources
Reasoning models persistently outperform non-reasoning models regardless of inference budget because training instills a reasoning protocol that makes additional tokens productive. The gap is fundamentally about deployment mechanisms and training structure, not raw capability.
LRMs don't break at complexity thresholds but at instance-novelty boundaries. Models fit instance-based patterns rather than generalizable algorithms, so any reasoning chain succeeds if trained on similar instances, regardless of length.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Evolving Deeper LLM Thinking
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Can Large Language Models Reason and Optimize Under Constraints?
- Reasoning Models Can Be Effective Without Thinking
- Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
- Branch-Solve-Merge Improves Large Language Model Evaluation and Generation