SYNTHESIS NOTE
Topics›Correct but Not Understood›this note

Why did Erdős problems become a popular AI testing ground?

Explores what makes Erdős problems attractive for evaluating large language models, including their mathematical domains, difficulty range, and the collaborative infrastructure that enables testing.

Synthesis note · 2026-10-06 · sourced from Correct but Not Understood

Kakaes reports that Erdős problems became a test bed for LLMs because, "by and large," they sit in number theory, combinatorics and graph theory, areas that have "proved more accessible than others to large language models," and because they "vary widely in difficulty." The excerpt's sharper point concerns what follows an answer. It keeps three things apart: a formal check, a model's re-check of another model's output, and a human's understanding. Aristotle, a tool from the startup Harmonic, was used to "certify that the proof held together logically" for Erdős 728. Bloom describes 100- to 200-page AI-generated papers that, in his words, "no human has read it, and no human is going to read it." Van Doorn, by contrast, says that when he reads an LLM's idea he will "digest it, try to understand it, simplify it, and generalize it."

The excerpt treats checking as an iterative loop, not a proof. Price asks a chatbot for a solution, feeds it "into a fresh instance of the chatbot, asking it to check the previous chatbot's work," and repeats "until he had what looked like a workable solution." Company harnesses automate the same loop, according to the excerpt. The word "looked" carries the argument: the procedure ends where another model stops objecting, which is a plausible answer rather than a proof someone has followed. The opposite failure shows in Erdős 333, where Barreto's claimed solution was already in a 1977 Erdős paper, a mistake he owned: "As someone who has fallen for this twice now, it's quite gut-wrenching."

Set against the neighbors, the heuristics note finds the same gap inside a model: transformers trained on orbital mechanics predict trajectories accurately yet apply task-specific laws instead of Newtonian ones, so accuracy alone does not show understanding. In the Erdős case the gap sits with the readers. The Darwin Gödel Machine note replaces formal proof with empirical validation on benchmarks; the Erdős cases keep a formal checker for logic and leave explanation to people. The AIDE2 note lists "untrustworthy wins" among practitioner problems, and the 333 episode is a concrete case of one. The excerpt also says most new results came from "hobbyists and undergraduates using publicly available LLMs" rather than corporate labs, and Bloom says such users are "not capable of verifying the output."

The excerpt measures none of this. It gives no count of AI results checked by experts, no error rate beyond the 333 case, and no verification detail for the DeepMind team's claim that "our most capable agent autonomously resolved 9 of 353 open Erdős problems." The OpenAI announcements of May 20 and August 1 are likewise lab claims, not checked findings. Barreto, Price and Lichtman appear without introduction, and the 333 passage starts mid-story, so passages appear to be missing. What the excerpt supports is narrower than the split it implies: formal certification and human comprehension are reported as different things, and comprehension looks like the part that lags. That lag is asserted, not measured, so a verified Erdős proof should not be read as understood until someone reports having read it.

Inquiring lines that read this note 11

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can we trust AI-generated mathematical proofs without understanding them?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 123 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Kakaes reports Erdős problems became a test bed for AI because they are accessible and vary widely in difficulty — and some AI proofs go unread