IMO gold in natural language: what changed between silver and gold
In 2024 AlphaProof needed the problems translated into Lean and days of compute to reach silver. This week Gemini Deep Think scored 35 of 42 working in English inside the 4.5 hour limit, and OpenAI reported the same score with an experimental model. What is known about the recipe, and why the Deep Think you can buy is not this one.
Two systems, one score
Google DeepMind announced on Monday that an advanced version of Gemini Deep Think scored 35 out of 42 at the 2025 International Mathematical Olympiad, solving five of the six problems within the 4.5 hour window that human contestants get. The proofs were graded by IMO coordinators, the same people who mark the students, and certified as gold standard. Two days earlier OpenAI had posted that an experimental research model scored the same 35, on the same five problems, with no tools and no internet.
Both systems missed problem 6. According to Simon Willison's write-up, only 6 of the 630 human competitors solved it fully, so that is the one place where the best students still outrun the models. The ordering of the announcements caused some noise. Demis Hassabis said DeepMind held its result until the IMO board's requested time, after verification and the student ceremony, and Noam Brown said OpenAI had checked with a board member and posted after the ceremony. The maths is the interesting part, so we will leave the etiquette there.
What silver required
Last year's result was 28 points, four problems solved, from AlphaProof and AlphaGeometry 2. That was a silver, and it was earned in a way that looks nothing like a student at a desk. The problems had to be translated by hand into Lean, the formal proof language, before the system could start. The search then ran for two to three days on some problems. The output was a machine-checked proof, which is a wonderful thing to have, but the pipeline had humans at the front and a calendar rather than a clock at the back.
Formal verification solved a real problem. If the proof compiles, it is correct, and the training signal is clean. The cost was that everything had to pass through a formal language that most mathematicians do not write and that has no ready expression for a great deal of olympiad reasoning. The 2024 system was strong at exactly the problems Lean made easy.
What gold required
The 2025 system reads the official problem statement in English and writes the proof in English. No translation, no checker. The reported recipe has three parts. The first is what DeepMind calls parallel thinking: instead of one long chain of reasoning, the model explores several candidate approaches at once and combines them before committing to an answer. The second is new reinforcement learning methods trained on multi-step reasoning and theorem-proving data, which is the long-horizon RL piece. The third is plain curation, a corpus of high-quality solutions plus what the post calls IMO-specific guidance, meaning hints about how to approach and write up olympiad problems.
The thing to notice is that the verification role moved. In 2024 Lean checked each step. In 2025 the model has to produce something a human grader accepts as rigorous, with no compiler in the loop. Whatever the RL setup was, it had to reward proofs that read as complete to a person, and the fact that IMO coordinators signed off on five of them says the reward held up on unseen problems.
The model you cannot use
DeepMind's post says a version of this model will go to trusted testers, including mathematicians, before any rollout to Google AI Ultra subscribers. OpenAI has described its model as experimental and not specific to the IMO. In both cases the thing that got gold is not the thing anyone outside the labs can run today. The Deep Think mode that ships in the Gemini app shares a name and presumably an ancestry, but the blog is careful to call the IMO entrant an advanced version.
That matters for how the result should be read. It is a demonstration of what a lab can do with a research model, a lot of parallel compute during the contest, and a curated training regime aimed at one competition. It is not yet a statement about what a product does at consumer prices. The gap between those two things is the part of the story we expect to close over the next year, and we would like to know the compute per problem when it does.
What we would want next
The result we want to see is the same system on problems written after its training cutoff, in a domain without olympiad-shaped training data, with the compute budget published. Five of six at the IMO in natural language is a strong signal that long-horizon RL on reasoning traces has crossed a threshold. Whether the threshold is about mathematics or about competition mathematics is the question the next twelve months should settle.
Sources
From the foundation