The weekend in order

The IMO announced its student results on Friday night, July 18. On Saturday morning OpenAI posted that an experimental reasoning model had solved five of the six problems for 35 out of 42 points, a gold-medal score, working from the natural-language problem statements. On Monday morning Google DeepMind announced that an advanced version of Gemini Deep Think had produced the same result, five of six, 35 of 42, within the 4.5 hour competition window, and that the IMO had officially graded and certified it. The IMO president, Gregor Dolinar, is quoted in the DeepMind post confirming the score.

Two systems, one score, and by Monday evening the argument was entirely about procedure. Demis Hassabis said DeepMind had not announced on Friday because it respected the IMO board's request that labs share results only after official verification by independent experts and after the students had received their acclamation. Noam Brown at OpenAI said they had spoken to an IMO board member who asked them to wait until after the award ceremony, and that they announced at around 1am Pacific after the ceremony concluded. An IMO coordinator quoted in the Hacker News thread said OpenAI had announced before the closing ceremony. Both of those things can be true depending on which ceremony is meant, and we do not think the timing is the interesting part.

Who graded what

The interesting part is the grading. OpenAI's proofs were scored by three former IMO medalists the company hired, who understood the official rubric and applied it. DeepMind's proofs were scored by the IMO's own coordinators, the people who grade the students, using the same criteria applied to the students, and the IMO put its name on the result. DeepMind had been working with the organisers since the previous year's silver-medal effort. OpenAI, according to Brown, was approached about a formal Lean track two months earlier and declined, and was never offered a natural-language option.

So we have one result that was graded by the body that defines the standard, and one that was graded by qualified people paid by the party being graded. In any other evaluation setting the second arrangement would be called self-report. That does not make the score wrong. Former medalists applying a published rubric to five proofs will probably reach the same numbers as the coordinators. But probably is doing the work in that sentence, and the whole point of an external grader is that you do not have to say probably.

What the certification covers and what it does not

DeepMind's post is careful about scope, and the care is instructive. The IMO's review confirmed that the submitted solutions were complete and correct. It did not validate the system, the process, or the underlying model. The coordinators saw proofs. They did not see how many attempts were generated before those proofs were submitted, how much compute was spent, or whether the 4.5 hour limit describes wall-clock time on a fixed budget or something looser. Those questions ran through the Hacker News thread and none of them are answered by the certification.

That is the honest boundary of what a third-party grader can do for a capability claim. It can tell you that the artefact meets the standard. It cannot tell you what it cost to produce the artefact or how often the system fails to produce one. For a benchmark that students take once, under one set of conditions, the artefact is the whole story. For a model, the selection process behind the artefact is most of the story, and neither lab has published it.

Why the protocol matters more than the score

Terence Tao's comment, relayed in the thread, makes the point from the other direction. Reformatted questions, extended time, tool access and team collaboration can each shift apparent capability by a large margin, and stacking several of them produces differences of orders of magnitude. The IMO's certification pins down some of those variables for the DeepMind run, in particular the time window and the problem statements. For the OpenAI run we have the company's description of its own conditions. Both may be accurate. Only one is checkable.

This is the same problem the evaluation community has been circling for two years, in a setting clear enough to see it plainly. A benchmark score means what the protocol says it means. When the lab controls the protocol, the grading, and the announcement, the score is a claim. When an independent body controls the grading and the lab publishes the protocol, the score becomes evidence. The 35 points are the same in both cases. What changed on Monday was the epistemic status of the number.

What we would want next

We would like the IMO, or a group it delegates to, to publish a standing protocol for AI entries before next year: problem delivery, time window, compute disclosure, number of attempts, and grading by coordinators, with results embargoed until the students have had their day. Labs that follow it get a certified score. Labs that do not can still post whatever they like, and the rest of us will know how to read it. Both labs solved the same five problems this year, and the sixth is a genuinely open question about what these systems cannot yet do. That is the conversation we would rather be having.

Sources

  1. TechCrunch, OpenAI and Google outdo the mathletes, but not each other
  2. Google DeepMind, Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the IMO
  3. Simon Willison, Gemini Deep Think and OpenAI at the IMO
  4. Hacker News discussion of the DeepMind announcement