Why the Legendary Erdős Problems Are Falling to AI
Source: Konstantin Kakaes, Quanta Magazine · 2026-08-03
On May 20, 2026, OpenAI made an announcement that shook the mathematical world. An internal AI model — one not available to the public — had come up with a counterexample to the “unit distance” problem, a conjecture made in 1946 by Paul Erdős, the prolific, itinerant Hungarian mathematician.
Erdős posed thousands of questions, but this one was special: It was both simple to explain and mathematically deep. It was the first historically significant proof to come from an AI model. Though the model’s result wasn’t definitive — human mathematicians would substantially improve on it within weeks — it was innovative, bringing in ideas from a distant branch of math that no one had successfully applied to this problem before. And it was influential: Within a few days, related techniques were used to solve other important problems.
Then on August 1, OpenAI announced that an unreleased model named Astra made 10 additional mathematical advances, including finding solutions to three more problems posed by Erdős.
Many mathematicians have hailed developments such as these as a phase transition in the mathematical capability of AI models. These models are “changing dramatically the way mathematical research is being done,” said Noga Alon of Princeton University, who has solved dozens of Erdős problems over his decades-long career.
But in all likelihood none of this would have happened had it not been for an English mathematician named Thomas Bloom.
Bloom has liked Erdős’ style for as long as he can remember. But he always found it hard to keep track of which problems had been solved and which had been forgotten entirely. So in early 2023, he decided to gather as many problems as he could into a list.
He intended it for his own use. But “I thought it would be easier if I could access it wherever I was,” he said; he figured he “might as well make a website, kind of with the expectation that maybe nobody would use it.” He gathered a couple hundred problems and launched erdosproblems.com. Bloom used ChatGPT to write the Python code that ran the website, which was, at the time, a remarkable thing for a large language model to be able to do. Using one to collaborate on the math itself still seemed like only a distant possibility.
His goal was not just to cross items off a list. He wondered if “modern day mathematics, often using techniques unknown by Erdős, could clear up many of these more obscure problems,” he wrote in a blog post. “We will then be left with a core of interesting, difficult problems, which can serve to demonstrate the limits of our knowledge.”
Then, in August 2025, some colleagues suggested that Bloom add a commenting function, so that people could talk about problems they were interested in. He was able to do so quickly, using ChatGPT to write the code. By now he’d cataloged nearly 1,000 problems.
Bloom’s timing was good. He made it possible for like-minded people to talk to one another, and that “really let a community build up,” he said. For the most part, comments were sporadic — a problem might attract a single comment pointing out an example or noting how hard the problem looked. But activity steadily grew, and some problems catalyzed nuanced mathematical discussions between strangers.
And so, in October 2025, van Doorn, now back at his day job, left the first comment on the page for Problem 1102. The problem, which Erdős posed in 1981, asks about properties of sets of “square-free” integers — that is, integers that have no repeated prime factors. (For instance, 30 is square-free because it is equal to 2 × 3 × 5, but 18 is not, because it is equal to 2 × 3 × 3; the 3 repeats.)
In early November, van Doorn shared progress toward an answer — which he’d figured out without relying on AI — as a comment on the problem page.
Bloom’s website, which has the look and feel of an earlier time, was becoming an example of the internet at its democratic best. “This entire collaboration would not have been possible without Tom’s website and the comments section there,” van Doorn said. It didn’t matter if you had tenure or not, if you were young or old, if you were at a fancy university or even at a university at all. If you wanted to work on math and had good ideas, you could find people to collaborate with.
But as the winter set in — around the same time that van Doorn found himself collaborating with Terry Tao — things started to change.
But a few hours later, another user pointed out that Erdős himself had provided a resolution to 333 in a paper published in 1977. Barreto owned up to the mistake. “My formal request to all members of the website is to put greater focus on literature search on the problems currently marked as open,” he wrote. “As someone who has fallen for this twice now, it’s quite gut-wrenching.”
Undeterred, he and Price kept at it, and by January 4, 2026, they’d used GPT-5.2 Pro to find a solution to Erdős 728, a problem about when certain numbers are divisible by other numbers. This time nobody could find prior work already proving it. Barreto used another AI tool called Aristotle (developed by a startup called Harmonic) to certify that the proof held together logically. Nat Sothanaphan, a software engineer and the only forum participant more prolific than Bloom, Tao, and van Doorn, had ChatGPT write up the formalized result and posted it online.
Price developed a methodology for how to ask LLMs to solve open questions. First, he would ask a chatbot for a solution. Then he would feed that solution into a fresh instance of the chatbot, asking it to check the previous chatbot’s work. He’d repeat this process until he had what looked like a workable solution. (This echoes some of the work that companies have been doing internally to create what they call harnesses or scaffolds, which automate the sort of iteration that Price does by hand.)
Barreto and Price’s papers represent just a fraction of the many Erdős problems solved at least in part by AI over the past few months. There are multiple reasons why these problems in particular have become such a fertile test bed for LLMs. The primary one is that, by and large, Erdős problems are in number theory, combinatorics, and graph theory, all areas of math that have proved more accessible than others to large language models. The problems also vary widely in difficulty and mathematical significance. This variation makes them appropriate for a nascent technology whose abilities also vary widely.
“A lot of my recent papers should be mostly credited to AI,” van Doorn said. “The ideas involved were ideas I did not come up with myself.” Like many people active on the Erdős site, van Doorn is excited about the way LLMs are allowing him to do more things more quickly. “If I read an idea by an LLM, I digest it, try to understand it, simplify it, and generalize it,” he said. He uses AI to better understand the math.
Not everyone holds themselves to this standard. “A big problem is AI is being used a lot by people who aren’t mathematicians, who don’t have a huge mathematical background and are not capable of verifying the output,” Bloom said. “They like to move fast, ask their AI to check it, it grows and grows. We’re seeing a lot more of these 100- to 200-page papers that people are posting. ‘I solved this theorem; I got AI to generate the proof and check the proof and write the paper.’ But no human has read it, and no human is going to read it. It’s a huge challenge now.”
Bloom was surprised that despite lots of attention from OpenAI, Google DeepMind, and several startups, most of the new results had come from hobbyists and undergraduates using publicly available LLMs, not from corporate labs using more advanced internal models.
But that would change a few weeks later, on May 20, 2026, when OpenAI announced that they had solved one of the most well known Erdős problems of all, the unit distance problem.
In the first months of 2026, the major tech companies began to see opportunity in erdosproblems.com. As Lichtman explained, “Erdős had over 1,000 papers. They were scattered.” An institute in Hungary had collected scanned images of many of the papers, but nobody had collected all the problems. “This kind of single repository that anyone can access — labs realized that this could effectively be a benchmark.”
In January, a team of 24 researchers led by Google DeepMind shared a paper solving four problems and finding old, forgotten solutions to nine more, after “using Gemini to systematically evaluate 700 conjectures labeled ‘Open’ in Bloom’s Erdős Problems database.” In May, a separate DeepMind team of 21 researchers announced that “our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars.” (As of this article’s publication, Bloom’s database contains 565 solved problems and 652 open ones, but the DeepMind team narrowed their search to problems that have been written in formal logic.)
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can we trust AI-generated mathematical proofs without understanding them?- Why did AI-generated proofs go unread by mathematicians?
- Did automated checking loops actually solve Erdős problems correctly?
- Why do AI systems excel at literature search but struggle with novel proofs?
- How does the Golod-Shafarevich criterion ensure infinitely many suitable number fields?
- Why do some Erdős problem solutions fail to resolve the originally intended claims?
- What evidence exists about whether AI-written proofs reduce mathematician learning?
- How does Kummer's theorem connect base-p digit carries to binomial coefficient divisibility?
- What pattern does this follow from OpenAI's earlier Erdős problem claim?
- Why did OpenAI's Erdős primality claim collapse under independent verification?
- What should mathematicians prioritize when machines can solve problems faster?
- Why do theorem provers crowd out other AI-for-mathematics approaches and tools?
- Can disclosure alone ensure independent verification of AI-assisted mathematical work?
- Can opaque AI tools suggest valid mathematics without external validation?
- How does search difficulty differ from construction difficulty in mathematics?
- Does verification by inspection scale for AI mathematics discoveries?
- Can pure mathematics provide an objective test that experimental science cannot?
- Can mathematical literature remain alive if no human experts understand it?
- Does automation always move the goalposts of what counts as real mathematics?
- What would it mean for mathematics to define itself before AI transformation?
- Why does the pigeonhole argument in CM fields produce unit-modulus points?
- What makes the transition from lattice points to planar distances work mathematically?