Why does mapping a grid from a bigger number system onto a flat plane create so many point pairs exactly one unit apart?
Why does projecting lattice vectors to a subfield increase shared unit distances?
This explores why the recent record-breaking unit-distance constructions take points from a high-dimensional number system and map them into the ordinary plane, and why that mapping produces so many pairs of points exactly one unit apart. The corpus covers the construction itself but does not spell out the projection step, so part of this answer fills in the reasoning.
This explores why mapping a lattice from a large number system down into the ordinary plane creates many pairs of points at exactly unit distance. The corpus covers this only partly. It describes the construction and names its key ingredient, but no note explains the projection step line by line. Read what follows as a guide to the two relevant notes. Where it goes beyond them, it says so.
The question comes from Erdős's unit distance problem: if you place n points in a plane, how many pairs can sit exactly one unit apart? Erdős's conjecture was that this count grows only slightly faster than n itself. What made OpenAI's unit distance counterexample succeed? reports that OpenAI found a counterexample using number fields, which are algebraic systems that extend the rational numbers. A follow-up made the gain explicit: How many unit distances can points in a plane have? shows a lattice construction that produces more than n^1.014 unit-distance pairs. That exponent is barely above 1, but it is enough to break the conjecture.
The notes don't spell out the intuition behind the projection, so here it is from standard number theory. A number field of high degree behaves like a space with many dimensions. The construction uses a CM field, which sits on top of a smaller real subfield. Viewing its elements through one complex embedding places them in the plane. The 'subfield' in the question is most likely that smaller real subfield. Two points end up one unit apart when the number connecting them has absolute value 1 in that embedding. Such numbers correspond to a norm condition relative to the subfield. A higher-dimensional field contains far more of these numbers than any small field can, and after projection each one becomes a unit-length step between lattice points. So the projection doesn't create distances. It collapses a large supply of algebraically distinct 'length-one' elements into one plane, where they all show up as unit distances.
The novel part, per What made OpenAI's unit distance counterexample succeed?, was letting the field's degree grow without limit. The danger with growing degree is that the arithmetic gets messy. The class number grows, roughly measuring how badly factorization breaks down, and so does the discriminant, which measures how spread out the field is. Golod–Shafarevich towers solve this. Together with a fixed prime that splits completely, they keep both quantities under control while the degree climbs. Controlling them is what lets many of those length-one elements be found among the points actually used. The tools are classical, from the 1960s. The new move was turning the growth knob.
There is also a mirror image in an unrelated part of the library. Do embedding dimensions fundamentally limit retrievable document combinations? shows that squeezing documents into a low-dimensional embedding limits which combinations of documents can be returned together. In search, the coincidences that projection forces are a failure. In the unit distance problem, those forced coincidences are exactly what the construction wants.
Sources 3 notes
OpenAI's counterexample to Erdős's unit distance conjecture builds on classical Golod–Shafarevich towers and number-field methods, but achieves its breakthrough by letting the degree [K : Q] → ∞. This allows a fixed split prime to suppress the class number and discriminant, enabling the construction of planar point sets with superlinear unit distances.
A lattice-based construction using CM fields and Golod-Shafarevich arguments produces sets of n points with more than n^1.014 pairs at unit distance, making explicit the exponent that a prior OpenAI result left unspecified. The bound is stated as 1.014114/C for an absolute constant C.
Communication complexity theory proves that for any embedding dimension d, there exists a maximum number of top-k document combinations that can be returned as results. Even embeddings optimized directly on test data hit this polynomial limit, demonstrated on trivially simple retrieval tasks.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Remarks on the disproof of the unit distance conjecture
- An explicit lower bound for the unit distance problem
- Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
- 'hello there the jacobian conjecture is false thanx': why a tiny social media post has mathematicians rethinking AI
- From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- On the Theoretical Limitations of Embedding-Based Retrieval
- Training for Compositional Sensitivity Reduces Dense Retrieval Generalization
- Why the Legendary Erdős Problems Are Falling to AI