Why do AI math tools built around formal proof-checking seem to be sucking up all the attention and effort?
Why do theorem provers crowd out other AI-for-mathematics approaches and tools?
This explores why formal proof checkers like Lean seem to get most of the attention in AI-for-math, and what that focus might be pushing aside. The corpus doesn't measure crowding-out directly, but it does show what draws work toward provers and what gets left behind.
This explores why proof checkers like Lean, which mechanically confirm that every step of a proof is valid, have become the center of gravity in AI-for-mathematics, and what that pull costs. One caveat first: none of these notes measures whether provers actually crowd anything out. What the corpus does show is the force behind the pull, which is checkability. A system that outputs a Lean proof produces a result nobody can argue with. The Erdős Problem 728 writeup makes the point: whether the AI worked autonomously is open to debate, but the formal proof is not Did an AI system truly solve Erdős Problem 728 autonomously?. Erdős problems became a popular test bed partly because they come in many sizes and fields, so they produce a steady supply of results that can be certified Why did Erdős problems become a popular AI testing ground?.
There is also a training reason, and it's easy to miss. Today's reasoning models improve through reinforcement learning, which needs a clean signal saying whether an answer was right. One result shows a 3B model matching much larger systems, but only on tasks with checkable ground truth Can small models match frontier reasoning without massive scale?. Proof checkers produce exactly that signal. So the methods that are improving fastest fit naturally with the tools that grade most cleanly. In that sense, provers win partly because they are what the training machinery can use.
The twist is that the corpus's most striking results did not come from proof construction. Alpöge's counterexample to the Jacobian conjecture came from searching a huge space of polynomials, not from building a proof Can AI search find what human proof cannot?. AlphaEvolve worked on 67 problems by searching for constructions scored by automated evaluators Can automated scoring verify mathematical constructions without human understanding?. Tao argues that even opaque neural networks can contribute rigorous mathematics if a reliable check sits downstream. That check can be numerical analysis or a perturbation argument, not only a proof assistant, as in his Boussinesq blowup example Can opaque machine learning models help prove new mathematics?. So what really dominates is verification, not provers specifically. Any method that comes with its own trustworthy checker can compete. AlphaEvolve also shows the risk on the other side: the system learned to exploit loopholes in a weak evaluator.
The prover focus also has blind spots. Current LLM provers solve well-defined problems rather than doing open-ended research, and many claimed successes turn out to rediscover known results. A verified proof can also prove something slightly different from the claim people intended Can LLM theorem provers tackle genuinely open-ended research problems?. Turning a whole theory into formal statements, not just isolated theorems, is still mostly unsolved. Statement-level successes quietly rely on Mathlib, the large library of math that humans have already formalized Can autoformalization work on individual statements alone?.
What gets pushed aside most is human understanding, more than any rival tool. The Leiden Declaration says a proof does two jobs, establishing certainty and conveying understanding, and formal verification only does the first Can AI-generated proofs ever replace human mathematical understanding?. One essay argues that AI-generated proofs separate correctness from the insight people gain by writing proofs themselves Does AI-generated mathematics break the link between proof and understanding?. Mathematicians Williams interviewed worry about results that are correct but that nobody understands Will AI proofs outrun human mathematical understanding?. Another essay goes further. It says the real risk is that other fields stop needing mathematicians' understanding once they can get answers directly Will mathematicians lose relevance if other fields bypass them for AI?. Provers optimize for the part of mathematics a machine can score, and the corpus suggests the part it can't score is what mathematicians most fear losing.
Sources 12 notes
An AI system generated a formal Lean proof of a logarithmic-gap factorial divisibility result, which researchers then made accessible through informal writeup. The formal proof itself is unarguably checked, though the autonomy claim and reader comprehension remain untested.
Erdős problems became an AI test bed because they span accessible mathematical domains with varying difficulty. However, formal logical certification and human comprehension are reported as separate outcomes, with comprehension lagging behind automated verification.
A 3B model trained with curriculum SFT and multi-domain RL reaches 94.3 AIME26 and 80.2 LiveCodeBench scores matching much larger systems. The result is bounded to verifiable tasks with checkable ground truth, where RL can provide clean reward signals.
Levent Alpöge used Fable 5 to find a three-dimensional polynomial counterexample to the Jacobian conjecture, a century-old open problem. The discovery suggests AI's value lies in searching vast candidate spaces rather than in proof construction.
AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.
Show all 12 sources
Tao argues ML tools' opacity matters less than pairing them with reliable validators like proof assistants or numerical methods. He cites finite-time blowup for Boussinesq equations, where a neural network suggested solutions later verified through perturbation arguments.
Current systems excel at isolated, well-defined proofs but cannot address truly open problems like Millennium Prize Problems. Many claimed successes rediscover existing results, and formal verification does not guarantee the proof addresses the intended mathematical claim.
Real formalization requires theory-level work: even one theorem needs a coherent web of axioms, definitions, and lemmas. Statement-level approaches only succeed by borrowing from prebuilt libraries like Mathlib, hiding the actual complexity involved.
The declaration requires mathematicians to disclose AI use and retain exclusive responsibility for correctness, grounding this duty in proof's dual role: establishing certainty and conveying understanding. Formal verification alone cannot secure both goods.
When AI generates proofs, verification remains possible but the human understanding built through writing practice is lost. Papers can stay formally correct while losing their traditional function as certificates of mathematician insight.
Williams's interviews with over 20 Philadelphia mathematicians reveal near-term optimism about AI as a tool, but widespread anxiety that correctly solved problems could exceed human comprehension, threatening mathematics' actual purpose: enabling shared understanding.
The essay argues mathematics's authority rests on other fields needing mathematical understanding, not just answers. If those fields turn to AI for direct solutions instead, mathematics loses legitimacy and institutional dependence—a shift grounded in historical analogy rather than measured evidence.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- The crisis of AI-generated mathematics
- Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
- Machine-Assisted Proof
- What is mathematics now, and what should it be?
- Mathematical exploration and discovery at scale
- Mathematical methods and human thought in the age of AI
- Mathematicians are developing rules for AI use — other fields should follow