INQUIRING LINE

A proof shows a result is true, and writing it can build understanding, so what happens when AI writes the proofs?

What evidence exists about whether AI-written proofs reduce mathematician learning?

This explores whether there's real evidence (not just worry) that leaning on AI to produce proofs erodes the understanding mathematicians normally build by working proofs out themselves.


This explores whether AI-written proofs actually reduce how much mathematicians learn, and the honest answer is that the corpus has no direct study of it. No paper here measures what mathematicians understand before and after using AI provers. What the collection does have is something almost as useful: a clear account of why people expect a loss, plus early field reports showing that understanding is already falling behind verification.

The main argument is that a proof does two jobs. It shows that a result is true, and the act of writing it builds the author's understanding. The essay on artificial mathematics argues that AI separates these jobs. Checking still works, but the understanding that came from the writing is gone, so a paper can be fully correct and no longer serve as evidence that anyone gained insight Does AI-generated mathematics break the link between proof and understanding?. The Leiden Declaration starts from the same two-job view and turns it into a rule. Humans must disclose AI use and stay solely responsible for correctness, because formal verification can secure truth but not understanding Can AI-generated proofs ever replace human mathematical understanding?. Both pieces are arguments and policy. Neither measures learning.

The closest thing to evidence comes from Erdős problems, the current proving ground for AI mathematics. Reporting on that work says that formal checking and human understanding are separate outcomes, and that some AI proofs are verified but not read closely by anyone Why did Erdős problems become a popular AI testing ground?. In the Erdős Problem 728 case, a Lean proof was machine-checked and then turned into ordinary mathematical prose. Whether human readers actually understood it was never tested Did an AI system truly solve Erdős Problem 728 autonomously?. AlphaEvolve shows the same pattern at larger scale. Automated scoring reliably confirmed constructions across 67 problems, but explaining why they work succeeded only some of the time Can automated scoring verify mathematical constructions without human understanding?. So the gap the essays describe is visible in practice. But a gap between verification and understanding is not the same as measured learning loss.

Terence Tao offers a different view. He treats AI as a source of suggestions and cares less about whether the tool is transparent than about pairing it with something that can check its output. In the Boussinesq blowup work, a neural network proposed candidate solutions, and humans then established them with their own perturbation arguments Can opaque machine learning models help prove new mathematics?. In that workflow the human still does the proving, so the learning may survive. This suggests the real variable is how AI gets used, not whether it is used at all.

One more comparison may sharpen your sense of the risk. Models face a version of the same problem. Fine-tuning a model on correct answers can raise its scores while reducing the information its reasoning steps actually add by 38.9%, so it lands on right answers through after-the-fact justification Does supervised fine-tuning improve reasoning or just answers?. Getting correct outputs without working through the steps that produce understanding is exactly what the mathematicians fear for themselves. The open research question is whether anyone will run the human version of that measurement.


Sources 7 notes

Does AI-generated mathematics break the link between proof and understanding?

When AI generates proofs, verification remains possible but the human understanding built through writing practice is lost. Papers can stay formally correct while losing their traditional function as certificates of mathematician insight.

Can AI-generated proofs ever replace human mathematical understanding?

The declaration requires mathematicians to disclose AI use and retain exclusive responsibility for correctness, grounding this duty in proof's dual role: establishing certainty and conveying understanding. Formal verification alone cannot secure both goods.

Why did Erdős problems become a popular AI testing ground?

Erdős problems became an AI test bed because they span accessible mathematical domains with varying difficulty. However, formal logical certification and human comprehension are reported as separate outcomes, with comprehension lagging behind automated verification.

Did an AI system truly solve Erdős Problem 728 autonomously?

An AI system generated a formal Lean proof of a logarithmic-gap factorial divisibility result, which researchers then made accessible through informal writeup. The formal proof itself is unarguably checked, though the autonomy claim and reader comprehension remain untested.

Can automated scoring verify mathematical constructions without human understanding?

AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.

Show all 7 sources
Can opaque machine learning models help prove new mathematics?

Tao argues ML tools' opacity matters less than pairing them with reliable validators like proof assistants or numerical methods. He cites finite-time blowup for Boussinesq equations, where a neural network suggested solutions later verified through perturbation arguments.

Does supervised fine-tuning improve reasoning or just answers?

Supervised fine-tuning improves final-answer accuracy on benchmarks but cuts Information Gain by 38.9 percent, meaning models generate correct answers through post-hoc rationalization rather than genuine inferential steps. Standard metrics miss this degradation because they only measure final correctness.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.