Leiden Declaration on Artificial Intelligence and Mathematics
Source: Leiden Declaration working group · 2026-06-02
Technological developments have repeatedly transformed the practice of mathematics. Recent artificial intelligence technologies, including symbolic and neural methods for the generation and formalization of mathematics, may already have initiated a significant chapter in this long history. Among researchers, artificial intelligence has produced a wide range of reactions: enthusiasm for its potential to yield new discoveries; intimidation by the pace of developments; indifference to these rapid changes; and concern for the implications, both for mathematics and in wider society.
Mathematicians have a choice about whether and how to adopt artificial intelligence in the conduct of their research. They also have a responsibility to ensure the continued flourishing of the discipline. This Declaration calls upon mathematicians to exercise this responsibility, and provides recommendations for individuals, institutions, government, and industry.
We base our recommendations on what we take to be characteristic values of mathematical research that we have a joint interest in preserving. Among these are the following:
There are many reasons to pursue mathematical research, ranging from intellectual curiosity to a desire to solve practical and societal problems. Underlying much of mathematics is the activity of proof. Mathematical proofs are regarded as conferring the highest degree of certainty to their conclusions, as well as imparting understanding of why their conclusions are true. These characteristics of proof support the scientific integrity of mathematics.
Results are attributable to specific authors who take credit for their discovery and assume responsibility for their correctness. These principles ground the merit-based standards to which we aspire in mathematical research.
Mathematical arguments are regarded as transparent and subject to independent verification. They may be extremely long or difficult, but in principle no proprietary knowledge or equipment should be required to understand them.
Mathematicians share a concern for proper evaluation of mathematical work relative to shared standards of depth, difficulty, and significance.
Mathematics produces not only a body of results, but also understanding, clarity, and judgment among the communities of mathematicians who have shaped them, often in the context of their own autonomously guided research. This expert knowledge is essential, both to effectively use mathematics, and to continue to articulate new and significant research questions. A key source of strength of the discipline has long been the autonomous shaping of the direction of research and the methods used to pursue it.
These characteristics of mathematics as a subject matter are also compatible with understanding mathematics as a human practice, and its place in the world. As mathematicians, and also as inhabitants of a shared world, we have a duty to care for other people and our environment.
Recent developments in artificial intelligence threaten each of these values, often in ways that disproportionately affect students and early-career mathematicians, and hence the long term future of the discipline.
Current automated techniques can produce plausible but unreliable (or even incorrect) arguments which are difficult to distinguish from correct mathematical proofs. This applies not only to informal arguments, but also to formalizations, where the difficulty lies in the translation between computer-encoded and human presentations of concepts. These fast-moving developments put our present system of review under increasing pressure, jeopardizing our ability to implement traditional standards for the correctness, transparency, and independent verifiability of proof.
Technologies that draw extensively on the published mathematical commons undermine the traditional system of attribution. Models trained on published works frequently return outputs that do not properly cite the human works they synthesize. Many current models are also built on data obtained by systematically exploiting licenses and access arrangements that were not made with artificial intelligence in mind, or indeed by simply violating copyright protections.
Technologies which affect the way in which mathematics is practiced may disturb the current system of incentives. The use of artificial intelligence — and thus also the sort of problems which it can address — may become incentivized for its own sake, disrupting our mechanisms for hiring, funding, and recognition. This disadvantages researchers who do not have access to the technologies or decision-making related to them, or who are unwilling to use technologies controlled by organizations whose values they do not share.
Proper evaluation is endangered if results are communicated through informal channels such as press releases or blog posts, often without any research paper or other disclosure of information necessary for scientific evaluation. This practice seeks publicity for new results on market timelines before the accepted processes of community evaluation in mathematics can take place. In many cases this leads to simplifications in reporting, such as overemphasizing the significance of automated tools and undervaluing the prior human contributions which have made those tools possible. Such oversimplification risks influencing public opinion in a way that not only damages perceptions of mathematics, but also misleadingly uses specific mathematical tasks as metrics for the general reasoning capacities of commercial products.
These developments put the autonomy of mathematics under threat. The increasing involvement of technology companies in mathematical research raises the risk that research questions may come to be prioritized because of their amenability to automated mathematics, rather than expert judgment of their deeper significance. Indeed, broader understanding of the field may be permanently lost in the process of automation. With university budgets under pressure, this reshaping also changes professional incentives in a manner which encourages the collaboration of researchers with technology companies on asymmetric terms. If left unchecked, these trends go beyond threatening researchers’ autonomy, affecting the scope and depth of mathematical research itself.
All of these challenges arise at a moment when the consequences of large-scale investment in artificial intelligence are being widely discussed in regard to warfare, mass surveillance, political disruption, and environmental damage. These raise grave ethical concerns. By failing to act, we run the risk of becoming complicit in the support of technologies which threaten much more than the practice of mathematics.
Transparently disclose the use of automated tools, including large language models, machine learning systems, proof assistants, and other mathematical software. Include a “Tool and computational resource disclosure” section in your papers; many journals, publishers, and professional organizations have already developed guidelines for this, and though the precise form of such a section will necessarily evolve, we encourage authors to live up to the spirit reflected in the UNESCO Recommendation on Open Science and the FAIR principles. When acting as a reviewer, abide by publisher guidelines. If the use of artificial intelligence is allowed, be transparent about how you used it, and take responsibility for any significant recommendations you make.
When automated techniques are employed in published mathematical research, the responsibility for the correctness and adequacy of the arguments and results, as well as for the completeness and accuracy of citations to relevant prior work, remains exclusively with the human authors.
Credit and responsibility continue to belong to humans within the mathematical community and should not be given to automated systems. Artificial intelligence may obscure, but does not replace, the collective human labor behind a result.
The known limitations of automated tools in properly attributing ideas create a corresponding obligation for proactive effort to find and credit the sources that made a new result possible. Where a satisfactory attribution is not possible, state this explicitly in the publication.
Mathematics has led to technology which greatly improves everyday life for many people, yet it also has applications in the development of technology for use in warfare, oppression, mass surveillance, and the undermining of democracy. Evaluate the ethical consequences of your research to the best of your abilities, and if necessary withdraw from harmful work. Only enter into external partnerships which respect the values articulated in this Declaration.
Professional organizations should keep abreast of technical developments and be proactive in making informed recommendations to members and to the broader community.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can we trust AI-generated mathematical proofs without understanding them?- What pattern does this follow from OpenAI's earlier Erdős problem claim?
- Can pure mathematics provide an objective test that experimental science cannot?
- What would it mean for mathematics to define itself before AI transformation?
- How does mathematical legitimacy depend on other fields needing mathematical understanding?
- Can scoring functions alone constitute verification of scientific discovery?
- What distinguishes empirical scoring from formal proof in discovery validation?
- Can formal verification certify a proof without human comprehension?
- Why did AI-generated proofs go unread by mathematicians?
- Did automated checking loops actually solve Erdős problems correctly?
- Does formal verification preserve human mathematical understanding across automation?
- How do plausible but incorrect AI arguments evade detection in mathematical proofs?
- Can disclosure alone ensure independent verification of AI-assisted mathematical work?
- What translation barriers exist between machine-encoded and human mathematical concepts?
- Why does the pigeonhole argument in CM fields produce unit-modulus points?
- Do proof assistants and neural networks fail in complementary ways?
- Can validated approximate solutions become exact mathematical proofs?
- Can a system recognize consequences of a theory without doing exact calculations?
- How does the Golod-Shafarevich criterion ensure infinitely many suitable number fields?
- Can proof assistants verify the full lattice construction argument formally?
- Why do some Erdős problem solutions fail to resolve the originally intended claims?
- What distinguishes rediscovering known results from genuine mathematical research?
- Can checking someone else's proof count as genuine mathematical understanding?
- What evidence exists about whether AI-written proofs reduce mathematician learning?
- Can a formally correct proof exist without the prover understanding the underlying mathematics?