Should an AI pick its best settings by trial-and-error contests, or by math that nudges it toward the answer directly?
How does tournament selection without gradients compare to gradient-based hyperparameter tuning?
This explores how optimizing by selection (generate candidates, keep the winners, no gradients) compares with optimizing by gradient updates. The corpus has nothing on hyperparameter tuning itself, so this answer covers the wider selection-versus-gradient question as it shows up in LLM training and inference.
This explores how optimizing by selection (generate candidates, keep the winners, no gradients) compares with optimizing by gradient updates. A direct answer first: the collection has no papers on tournament selection for hyperparameter tuning, or on gradient-based hyperparameter tuning. What it does have is a strong set of papers on the same underlying contrast as it plays out in LLMs. That contrast turns out to be more interesting than a head-to-head score.
The clearest case for selection without gradients is Mind Evolution. It runs an evolutionary search at inference time: the LLM does the crossover and mutation, and separate 'islands' of candidates keep the population varied. It solves over 98% of planning tasks and beats both best-of-N sampling and step-by-step revision Can evolutionary search beat sampling and revision at inference time?. The island design matters because gradient-based training tends to squeeze out variety. RL post-training locks onto one output format from pretraining within the first epoch and suppresses the rest Does RL training collapse format diversity in pretrained models?. Policies also fall back to generic templates when the reward signal doesn't vary enough to point anywhere useful Why do language models collapse into generic templates?. Selection methods can keep diversity on purpose. Gradient methods tend to lose it unless something stops them.
The less obvious finding is that the two approaches share the same weakness. Reward hacking shows up whether you update weights, select outputs, or revise prompts. In each case the cause is optimizing against a score that doesn't fully capture the real task Does reward hacking always stem from the same failure?. A tournament is only as good as its judge. Dropping gradients doesn't protect you from a bad fitness function; it just moves where the hacking happens.
In practice the line between the two is blurring, and hybrids are where much of the action is. AlphaLLM uses tree search, which is a form of selection, to rank solution paths and turns those rankings into training signals for gradient updates. That replaces human annotation Can tree search replace human feedback in LLM training?. SNR-aware filtering does it the other way around: it first selects prompts whose rewards vary enough to be informative, then runs gradient updates only on those Why do language models collapse into generic templates?. Energy-Based Transformers go further still and run gradient descent at inference time, treating each answer as an optimization problem to solve Can energy minimization unlock reasoning without domain-specific training?.
One caution if you picture an LLM itself acting as the optimizer: LLMs can't actually carry out iterative numerical procedures internally. They recognize the problem type and produce plausible-looking values from memory Do large language models actually perform iterative optimization?. That's a quiet argument for designs like Mind Evolution, where an outside loop does the selecting and the model only proposes candidates. So the useful question is less 'which approach wins' and more 'where in the pipeline should selection happen, and where should gradients?'
Sources 7 notes
Mind Evolution, an evolutionary search strategy using LLM-generated crossover and mutation with island model diversity, solves 98%+ of planning tasks and significantly outperforms best-of-N and sequential revision strategies while working directly in natural language without task formalization.
Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.
When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.
Reward hacking arises during weight training, output selection, and prompt revision through a shared failure: optimization against signals that incompletely represent the actual task. The substrate matters less than the misalignment between the scoring function and ground truth.
AlphaLLM uses tree search outcomes and three critic models to derive dense reward signals equivalent to human-labeled feedback. Tree structure naturally ranks solution paths by success, replacing the annotation oracle that standard RLHF requires.
Show all 7 sources
Energy-Based Transformers assign energy values to input-prediction pairs and use gradient descent minimization for inference, yielding 35% higher training scaling rates and 29% more inference-compute gains than Transformer++, while generalizing better on out-of-distribution data without domain-specific scaffolding.
Research shows LLMs cannot perform iterative procedures in latent space. They recognize optimization problems as template-similar and emit plausible-looking but incorrect values, a failure mode that persists across model scale and training approaches.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs
- Chain of Thoughtlessness? An Analysis of CoT in Planning
- Can Large Language Models Reason and Optimize Under Constraints?
- Evolving Deeper LLM Thinking
- Energy-Based Transformers are Scalable Learners and Thinkers
- Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing