What made OpenAI's unit distance counterexample succeed?
Researchers trace OpenAI's refutation of Erdős's unit distance conjecture to classical number theory tools, but pinpoint one novel ingredient: letting field degree grow without bound. Why did this shift unlock a solution that defeated many human attempts?
The remarks present a "short, digested, human-verified version" of an OpenAI internal model's counterexample to Erdős's unit distance conjecture, the conjecture that n points in the plane determine at most n^{1+o(1)} unit distances. Their Theorem 1.1 gives point sets with at least |P_i|^{1+ε} unit distances for some fixed ε > 0, which refutes the conjecture. The abstract says the argument "relies crucially on ideas that may, at least in retrospect, be attributed to" Ellenberg–Venkatesh, Golod–Shafarevich, and Hajir–Maire–Ramakrishna. The remarks locate the novelty more narrowly: the construction of Golod–Shafarevich towers with infinitely many split primes "already appears in the literature", and "a novel ingredient of the AI argument is to take [K : Q] → ∞."
The mechanism runs through two lemmas. Lemma 2.1 turns a lattice with many unit-modulus points in its polydisc into a planar point set with many unit distances, provided the count of those points grows fast enough relative to the lattice's skewness. Lemma 2.2 supplies the unit-modulus points by a pigeonhole argument in a CM field. Golod–Shafarevich towers keep the root discriminant bounded as [K : Q] → ∞, so a fixed split prime "drowns out the main enemies, the class number h(K) and discriminant Disc K." The AI's chain of thought, quoted in the introduction, points the same way: "Maybe that enormous degree is not just an annoyance but a source of possible counterexamples. Number fields deserve a closer look."
The excerpt keeps two operations apart. The construction is "human-verified", so the authors checked the AI's argument; the remarks are also "digested", a somewhat simplified and generalized version a human can follow. The excerpt does not say how well humans followed the original AI proof, and its own need for simplification is the only evidence about that. Where the reflections turn to understanding, they concern why the search succeeded. Noga Alon writes that "the AI was able to do here what lots of excellent human researchers tried and failed to do," and holds this "with or without a full agreement with these reasons" that colleagues have offered. Thomas Bloom calls the result "both surprising and impressive." On this excerpt, the gap between correct and understood sits in the explanation of the success more than in the proof.
Set against the nearest notes, this is a different case from Can identical outputs hide broken internal representations?: there a benchmark can hide broken internals, while here a checkable proof confirms correctness without explaining why the search worked. It shares with Do foundation models learn world models or task-specific shortcuts? the pattern of right outputs without the general mechanism, but the output here is an argument people can audit. It also differs in what counts as validation: the Can AI systems improve themselves through trial and error? note describes replacing proofs with benchmarks, and this result returns to a human-checked proof.
The excerpt does not establish several things. It gives no value for ε in Theorem 1.1, and Lemma 2.1 leaves u, v, and δ unquantified beyond the condition u > 36v/π, so the power-law gap is shown in principle, not measured. It omits the AI's transcript and describes no verification procedure beyond the authors' own label "human-verified." It also leaves open the explanation Alon himself hedges. At the strength the evidence allows, the implication is narrow: a human-checkable argument for a fixed power-law gap exists in these sections, its new move is the growing field degree, and whether this kind of AI success generalizes to other open problems is a separate question the source does not answer.
Inquiring lines that read this note 10
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can we trust AI-generated mathematical proofs without understanding them?- Why does the pigeonhole argument in CM fields produce unit-modulus points?
- What makes the transition from lattice points to planar distances work mathematically?
- How does the Golod-Shafarevich criterion ensure infinitely many suitable number fields?
- How close is the n^1.014 bound to the known upper bound of n^4/3?
- Why do some Erdős problem solutions fail to resolve the originally intended claims?
- Why did the Jacobian conjecture resist proof for over a century?
- What pattern does this follow from OpenAI's earlier Erdős problem claim?
- Why did OpenAI's Erdős primality claim collapse under independent verification?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can identical outputs hide broken internal representations?
Can neural networks produce correct outputs while having fundamentally fractured internal structure that prevents generalization and creativity? This challenges our assumptions about what performance benchmarks actually measure.
contrast: there performance hides broken internals; here a checkable proof confirms correctness without explaining why the search worked.
-
Do foundation models learn world models or task-specific shortcuts?
When transformer models predict sequences accurately, are they building genuine world models that capture underlying physics and logic? Or are they exploiting narrow patterns that fail under distribution shift?
parallel: correct outputs without the general mechanism, though here the output is a proof people can audit.
-
Can AI systems improve themselves through trial and error?
Explores whether replacing formal proof requirements with empirical benchmark testing enables AI systems to successfully modify and improve their own code iteratively, and what mechanisms prevent compounding failures.
contrast on validation: that note replaces proofs with benchmarks, while this result rests on a human-checked proof.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Remarks on the disproof of the unit distance conjecture
- An explicit lower bound for the unit distance problem
- Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
- Why the Legendary Erdős Problems Are Falling to AI
- 'hello there the jacobian conjecture is false thanx': why a tiny social media post has mathematicians rethinking AI
- From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- Mathematicians are developing rules for AI use — other fields should follow
- The crisis of AI-generated mathematics
Original note title
OpenAI's unit distance counterexample is built from classical number-field tools — letting the field degree grow is the novel ingredient