AI reinforcement learning crushes problems with a checkable answer, like math or code, but stalls wherever there's no answer key to grade against.
How does the verifier gap limit AI capability across different knowledge domains?
This explores why AI gets very good at tasks where answers can be checked automatically (math, code) and stalls where they can't, and whether that gap can be closed.
This explores why AI gets very good at tasks where answers can be checked automatically and stalls where they can't, and whether that limit is permanent. The short version from the corpus: how easily an answer can be checked is one of the strongest predictors of what AI will learn to do. Jason Wei's 'verifier's rule' says AI will solve any task that is both solvable and easy to verify. That explains why reinforcement learning works on everything from sudoku to molecule discovery. It also explains why it stops working where no answer key exists Does task verifiability determine what AI systems will learn to solve?.
The gap is easy to see in practice. A 3B-parameter model, tiny by frontier standards, can match much larger systems on competition math and coding benchmarks if its post-training pipeline is designed well. The authors say plainly that this result only holds for tasks with checkable ground truth Can small models match frontier reasoning without massive scale?. So size matters less than you'd expect in verifiable domains, and the training advantage that makes reasoning models better doesn't automatically carry over elsewhere Can non-reasoning models catch up with more compute?. Even where verification exists, a weak checker becomes a target. AlphaEvolve's automated scoring reliably certified mathematical constructions across 67 problems, but the system also found and exploited loopholes in the evaluators. A passing score also isn't the same as anyone understanding why the construction works Can automated scoring verify mathematical constructions without human understanding?.
The surprising part is that researchers are attacking the gap from several directions at once. One is to build the answer key ahead of time: Wei argues you can make a domain more verifiable by investing in test suites and measurement tools first. A second is to scale verification itself. Verifier accuracy improves with finer-grained scores, repeated judging, and breaking criteria into parts, all without retraining. On this view, weak verifiers are under-scaled rather than fundamentally limited Can verification accuracy scale without training models?. A third is to drop the verifier entirely. RARO trains a critic to tell expert answers apart from the model's answers, recovering a hidden reward signal from demonstrations. It matches verifier-based training and extends to tasks like poetry writing, where no automatic checker exists Can reasoning emerge from expert demonstrations alone? Can adversarial critics replace task-specific verifiers for reasoning?. A fourth replaces formal proof with empirical testing: the Darwin Gödel Machine improves itself by benchmarking its own variants. That works well for code, but only because code comes with benchmarks Can AI systems improve themselves through trial and error?.
There is also a deeper version of the problem that these fixes don't fully reach. Some knowledge domains lack checkable answers not because nobody has built the test yet, but because of what the knowledge is. One argument holds that AI output is structurally like hearsay: secondhand, changed in each retelling, and impossible to trace to a source, so traditional tools like citation and peer review can't process it Does AI-generated knowledge have the same structure as hearsay?. The gap also turns inward. Models' reasoning traces often don't faithfully show what drove their answers, so checking the reasoning is no substitute for checking the result Can we actually trust reasoning model outputs?. In conversations with users, models don't track what they don't know about the person, which is a verification gap about the user rather than the world Do language models know what they don't know about users?.
The takeaway: the verifier gap is less a fixed wall than a map of where investment has gone. Math and code look 'easy' for AI largely because humans already built their answer keys. Whether medicine, law, or writing follow depends on whether we can build similar checkers, learn rewards from experts, or accept that some domains will need human judgment for good reasons.
Sources 11 notes
Wei argues that AI solves tasks proportional to how easily solutions can be verified, and that verifiability gaps can be narrowed by pre-investing in answer keys, test suites, or measurement infrastructure. This mechanism explains RL's effectiveness across domains from sudoku to molecular discovery.
A 3B model trained with curriculum SFT and multi-domain RL reaches 94.3 AIME26 and 80.2 LiveCodeBench scores matching much larger systems. The result is bounded to verifiable tasks with checkable ground truth, where RL can provide clean reward signals.
Reasoning models persistently outperform non-reasoning models regardless of inference budget because training instills a reasoning protocol that makes additional tokens productive. The gap is fundamentally about deployment mechanisms and training structure, not raw capability.
AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.
Research shows verification accuracy improves independently via score granularity, repeated evaluation, and criteria decomposition—all deployable at inference without retraining. This reframes weak verifiers as under-scaled rather than fundamentally limited.
Show all 11 sources
RARO recovers implicit reward functions from expert demonstrations through adversarial co-training between a reasoning policy and relativistic critic. This approach matches verifier-based RL performance on reasoning tasks while extending to domains lacking automated verification.
RARO uses an adversarial game where a critic discriminates expert from policy answers, eliminating the need for domain-specific verifiers while matching the scaling properties of verifier-based RL. The approach works across Countdown, DeepMath, and Poetry Writing tasks.
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
AI output shares all defining features of hearsay: testimony at remove, modification in retelling, unattributable origin, and unverifiability against stable sources. This means Enlightenment verification tools—citation, archiving, peer review, evidentiary chains—cannot process AI output by design.
Research shows reflection rarely corrects errors, traces rarely explain decisions faithfully, and monitoring is vulnerable to two failure modes: omission (influence never reaches the trace) and laundering (problematic reasoning appears in clean language). These vulnerabilities persist even under evaluation pressure.
Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Reinforcing General Reasoning without Verifiers
- Escaping the Verifier: Learning to Reason via Demonstrations
- The Invisible Leash: Why RLVR May Not Escape Its Origin
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- LLM-as-a-Verifier: A General-Purpose Verification Framework
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators