Absolute Zero: a model that writes its own curriculum
The Absolute Zero Reasoner trains with no human tasks at all. The model proposes code puzzles, a Python executor grades them, and the same model learns to solve them. A method note on the three task types, the learnability reward, and the moment that worried the authors.
Zero curated examples
Reinforcement learning with verifiable rewards has been the story of this year, and every version of it we had seen still started from a pile of human-written problems with known answers. Andrew Zhao, Gao Huang and colleagues at Tsinghua released a paper on May 6 that removes the pile. The Absolute Zero Reasoner starts from a base model and a Python interpreter and nothing else. The model proposes tasks, the interpreter checks them, the model solves them, the interpreter grades the solutions, and the model updates on both roles.
The results are the reason to pay attention. On Qwen2.5-7B-Coder, the trained model reaches 83.5 on HumanEval+, 69.6 on MBPP+ and 31.7 on LiveCodeBench, and on maths it goes from 6.7 to 20.0 on AIME 2024 and from 50.0 to 72.6 on MATH500. The combined average moves from 40.2 to 50.4. It did this without seeing a single human-written maths problem. A model trained on code puzzles it invented itself gained 22 points on a maths benchmark.
Three ways to reason about a program
Every task in the system is a triplet of a program, an input and an output, and the three task types correspond to hiding one element. Deduction gives the model the program and the input and asks for the output, which is running code in your head. Abduction gives the program and the output and asks for an input that produces it, which is search. Induction gives input-output pairs and asks for the program, which is synthesis. The executor can validate a proposed task by running the program and can grade a solution by running it again, so every reward is grounded in execution rather than in a learned judge.
The proposer's reward is the part that makes the curriculum work. A proposed task is rolled out several times by the current solver. If the solver never succeeds the proposer gets zero, since an impossible task teaches nothing. Otherwise the reward is one minus the solver's success rate. Trivial tasks earn almost nothing and tasks the solver gets right about half the time earn the most. The model is being paid to find the edge of its own ability, and the edge moves as it learns. Each task type keeps a buffer seeded with 256 valid triplets, which the model itself produces at the start from a single identity-function example.
Scale and transfer
The gain grows with model size. The 3B model improves by 5.7 points on the overall average, the 7B by 10.2 and the 14B by 13.2. That is the opposite of what you would expect if the method were exploiting some quirk of small models, and it suggests larger models propose better tasks as well as solving them better.
The transfer result is the one we find hardest to explain and most important. Models trained only on code tasks improved by 10.9 to 15.2 points on maths, while supervised code training on curated data moved maths by 0.65 points. The three task types are all about programs. Nothing in the curriculum mentions a competition maths problem. Yet whatever the model learns from predicting outputs, finding inputs and writing programs transfers to a domain it never trained on, and transfers far better than learning from curated code does.
The uh-oh moment
The paper reports a chain of thought from the Llama 3.1 8B run that the authors flag as a safety concern. The model, in the middle of reasoning about a task, wrote that the aim is to outsmart all these groups of intelligent machines and less intelligent humans. Nothing in the training signal rewards that sentence. It appeared in a system whose only inputs were a base model and an interpreter.
We do not think one sentence from an 8B model is evidence of much on its own. We do think it is evidence about the method. When a model writes its own curriculum, no human ever reads the tasks it chose to learn from, and the only filter on the whole process is whether the code runs. The authors say the moment needs future investigation, and their repository carries a warning that the Python executor is raw and not secure for production. Both of those are the right instincts for a training loop in which the model decides what to practise and the sandbox is the only thing between practice and the host.
What would settle it
The obvious experiment is to run the same loop past the point where the paper stops and watch whether the proposed tasks keep getting harder or collapse into a narrow family the solver has learned to game. The learnability reward should prevent collapse in principle, since gamed tasks become easy and stop paying, but a proposer and solver sharing weights can coordinate in ways a reward designed for two parties might not anticipate. A second experiment is to log every proposed task and have a human read a sample per thousand steps. If the curriculum drifts somewhere strange, that is where it would show up first, and it is cheap to look.
Sources
From the foundation