What does correctness of outputs tell us about reasoning?
IMO graders verified that Gemini's proofs were mathematically correct, but their review excluded the model's processes and training. This raises whether certified right answers demonstrate genuine understanding or only output accuracy.
Google DeepMind claims that an advanced version of Gemini Deep Think solved five of the six problems at the 2025 International Mathematical Olympiad, earning 35 of 42 points, a gold-medal score. The result is the lab's own claim, but the excerpt says the answers were graded by IMO coordinators "using the same criteria as for student solutions," and that the IMO "confirmed that our submitted answers are complete and correct solutions." That is the verified part: the proofs were checked by the competition's own graders. The only comprehension-like statement is that the graders found the solutions "clear, precise and most of them easy to follow." That describes readability. The excerpt does not say the reasoning was explained or understood in human terms.
The excerpt credits the result to a reasoning mode that "simultaneously explore[s] and combine[s] multiple possible solutions before giving a final answer, rather than pursuing a single, linear chain of thought." The model was also trained with novel reinforcement learning on multi-step reasoning and theorem-proving data, given a curated corpus of solutions, and given general hints in its instructions. The excerpt contrasts this with 2024, when AlphaGeometry and AlphaProof needed problems translated into Lean and two to three days of computation. This year the model worked end-to-end in natural language within the 4.5-hour limit.
The parallel-exploration mechanism is a different axis from the result in Does more thinking time always improve reasoning accuracy?, which concerns what happens when a single chain is made longer. The excerpt attributes its gains to breadth of exploration, not length, and gives no measurement that would test the non-monotonic finding. The certification also bears on Do automated benchmarks hide what frontier AI systems can really do?. An expert-graded competition is a stronger check than an automatically graded benchmark, but it still stops at the output. The accuracy of a result says little about how it was reached, which is the gap Can we measure how deeply a model actually reasons? tries to close from inside the model.
The excerpt does not establish how the system reached its proofs, and the lab says so itself: the IMO's "review does not extend to validating our system, processes, or underlying model." Five correct proofs on one competition's problems show that the outputs were correct on that occasion. They do not show that the training data, the hints or the parallel procedure produced the capability, or that it transfers beyond olympiad problems. At the strength the evidence allows, "correct" holds for these proofs, while "understood" does not yet hold for the system that produced them, and the public claim should be read at that strength.
Inquiring lines that read this note 14
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can we trust AI-generated mathematical proofs without understanding them?- Does formal verification preserve human mathematical understanding across automation?
- Can validated approximate solutions become exact mathematical proofs?
- What distinguishes rediscovering known results from genuine mathematical research?
- Can checking someone else's proof count as genuine mathematical understanding?
- Can a formally correct proof exist without the prover understanding the underlying mathematics?
- What makes a Lean proof an unarguable check compared to other mathematical verification methods?
- Why did OpenAI's Erdős primality claim collapse under independent verification?
- Can mathematics remain trustworthy when results bypass peer review entirely?
- Does publishing proofs without showing the verification process undermine mathematics?
- What verification methods can prove AI mathematical proofs are sound?
- Can pure mathematics provide an objective test that experimental science cannot?
- Does a correct proof preserve mathematical value without human comprehension?
- How does mathematical legitimacy depend on other fields needing mathematical understanding?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do automated benchmarks hide what frontier AI systems can really do?
Benchmarks optimize for auto-gradable, short, cheap tasks. But real AI capability emerges in long-horizon, messy, open-ended work. How much capability are we missing—or wrongly inflating—by relying on benchmark scores alone?
extends the evaluation question: expert certification checks outputs, leaving the system unassessed
-
Does more thinking time always improve reasoning accuracy?
Explores whether extending a model's thinking tokens linearly improves performance, or if there's a point beyond which additional reasoning becomes counterproductive.
contrasts: the source credits parallel exploration of several solutions, not a longer single chain
-
Can we measure how deeply a model actually reasons?
What if reasoning quality isn't about length or confidence, but about how much a model's predictions shift across its internal layers? Can tracking these shifts reveal genuine thinking versus pattern-matching?
output-level correctness says nothing about internal effort; this note is one attempt to measure it
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad
- Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
- Models May Behave Worse When Eval Aware
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- AI companies must work with the research community to protect attribution
- Reinforcing General Reasoning without Verifiers
Original note title
Google DeepMind reports IMO graders certified the Gemini Deep Think proofs as correct without validating the system behind them