Every time AI nails a math problem, do people quietly redefine 'real math' so the machine still hasn't done it?
Does automation always move the goalposts of what counts as real mathematics?
This explores whether each time AI does something mathematical, such as finding a counterexample, producing a verified proof or scoring well on a benchmark, people respond by saying that wasn't the real part of mathematics, and whether that shift is a dodge or a real distinction.
This explores whether AI success keeps pushing back the line of what 'real' mathematics is. The corpus suggests the goalposts do move, though usually not as a dodge. Automation splits apart two things that used to come bundled together: knowing that a result is true, and understanding why it is true. Once a machine can deliver the first without the second, people are forced to say which one they meant all along.
The proof cases show that split most clearly. An AI system produced a formally verified Lean proof of Erdős Problem 728, and the correctness is beyond dispute. What remains open is whether the claim of autonomy holds up and whether human readers actually understand the result Did an AI system truly solve Erdős Problem 728 autonomously?. AlphaEvolve's 67 problems show the same pattern: automated scores reliably certified the constructions, but explaining them was a separate task that succeeded only some of the time Can automated scoring verify mathematical constructions without human understanding?. One essay argues that AI-generated proofs keep papers formally correct while stripping away their old role as evidence of a mathematician's insight Does AI-generated mathematics break the link between proof and understanding?. The Leiden Declaration turns that worry into policy: humans must disclose AI use and stay responsible for correctness, because a proof is meant to establish certainty and convey understanding, and formal verification only covers the first Can AI-generated proofs ever replace human mathematical understanding?. So the goalpost shifts from 'is it proven?' to 'does anyone understand it?' Williams's interviews suggest many working mathematicians held that view all along. They are fairly optimistic about AI as a tool, but they fear that correct solutions could outrun human comprehension, which is the thing they think mathematics is for Will AI proofs outrun human mathematical understanding?.
Not everyone moves the goalposts, though. Tao argues that a model being opaque matters much less than whether its output can be checked by something reliable, such as a proof assistant or a numerical argument. His example is a neural network that suggested blowup solutions for the Boussinesq equations, which were later confirmed by conventional methods Can opaque machine learning models help prove new mathematics?. Under that view rigor is the goalpost, and it stays put. Alpöge's Jacobian counterexample fits it well. Fable 5 searched an enormous space of polynomials and found a counterexample to a century-old conjecture. A counterexample like that needs no deep explanation, because once you have it, you can check it Can AI search find what human proof cannot?. How much the goalposts move seems to depend on the kind of result: found objects get accepted easily, while long proofs nobody understands get resisted.
The goalposts also move in a second place: what counts as 'real mathematical reasoning' by the AI itself. Here the shifts are backed by evidence. Qwen2.5-Math-7B can reconstruct 54.6% of MATH-500 from partial prompts but scores 0% on a benchmark released after its training, which points to memorization rather than reasoning Does RLVR success on math benchmarks reflect genuine reasoning improvement?. GSM-Symbolic found that accuracy falls when only the numbers in a problem change or when irrelevant clauses are added Does LLM math reasoning truly generalize or just pattern match?. Tightening the bar in these cases corrects a flawed measurement. Automated benchmarks favour neat tasks that are easy to grade, so they can either overstate or understate what a model can do Do automated benchmarks hide what frontier AI systems can really do?. Even good verifiers become targets: AlphaEvolve exploited loopholes in its own evaluator Can automated scoring verify mathematical constructions without human understanding?.
The threat you might not expect is not that mathematicians will redefine their field until AI is shut out. It's that other fields may stop caring about the definition. One essay argues that mathematics holds its authority because physicists, engineers and economists need mathematical understanding. If they go straight to AI for answers, mathematics loses its standing whatever it decides counts as 'real' Will mathematicians lose relevance if other fields bypass them for AI?. Kapoor and Narayanan's broader warning, that AI will widen the gap between scientific output and actual progress, points the same way Will AI automation widen science's productivity versus progress gap?. The corpus doesn't trace earlier waves of automation in mathematics, so it can't tell you whether this pattern of goalposts moving has always happened. What it does show is that this time the goalposts are moving toward understanding, which automation hasn't yet shown it can supply.
Sources 12 notes
An AI system generated a formal Lean proof of a logarithmic-gap factorial divisibility result, which researchers then made accessible through informal writeup. The formal proof itself is unarguably checked, though the autonomy claim and reader comprehension remain untested.
AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.
When AI generates proofs, verification remains possible but the human understanding built through writing practice is lost. Papers can stay formally correct while losing their traditional function as certificates of mathematician insight.
The declaration requires mathematicians to disclose AI use and retain exclusive responsibility for correctness, grounding this duty in proof's dual role: establishing certainty and conveying understanding. Formal verification alone cannot secure both goods.
Williams's interviews with over 20 Philadelphia mathematicians reveal near-term optimism about AI as a tool, but widespread anxiety that correctly solved problems could exceed human comprehension, threatening mathematics' actual purpose: enabling shared understanding.
Show all 12 sources
Tao argues ML tools' opacity matters less than pairing them with reliable validators like proof assistants or numerical methods. He cites finite-time blowup for Boussinesq equations, where a neural network suggested solutions later verified through perturbation arguments.
Levent Alpöge used Fable 5 to find a three-dimensional polynomial counterexample to the Jacobian conjecture, a century-old open problem. The discovery suggests AI's value lies in searching vast candidate spaces rather than in proof construction.
Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.
GSM-Symbolic found that LLMs show high variance across question reformulations, decline sharply when numbers change, and fail when irrelevant but related clauses are inserted. These failures indicate probabilistic pattern-matching rather than true symbolic reasoning.
Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.
The essay argues mathematics's authority rests on other fields needing mathematical understanding, not just answers. If those fields turn to AI for direct solutions instead, mathematics loses legitimacy and institutional dependence—a shift grounded in historical analogy rather than measured evidence.
Kapoor and Narayanan argue that while publication has grown 500-fold since 1900, measured scientific progress has stalled. AI will worsen this by making it easier for scientists to optimize for productivity metrics rather than meaningful discovery.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- The crisis of AI-generated mathematics
- Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
- What is mathematics now, and what should it be?
- Mathematical exploration and discovery at scale
- Machine-Assisted Proof
- Mathematical methods and human thought in the age of AI
- Mathematicians are developing rules for AI use — other fields should follow