Textbooks are all you need: the small model, synthetic data bet
A 1.3B parameter model trained on under 7B tokens just scored 50.6% on HumanEval. What the phi-1 paper actually shows, and what it leaves open.
What Microsoft trained
A 1.3 billion parameter code model called phi-1, trained for four days on eight A100s, reached 50.6% pass@1 on HumanEval and 55.5% on MBPP. The paper from Suriya Gunasekar, Sebastien Bubeck, Yuanzhi Li and sixteen coauthors at Microsoft Research went up on arXiv on June 20 and the title, Textbooks Are All You Need, tells you exactly what they think the result means.
The training set is small by current standards. The whole thing is under 7B tokens. About 6B of that is code filtered from The Stack and StackOverflow for what the authors call textbook quality, and about 1B is synthetic textbooks generated with GPT-3.5. Then there is a finetuning set, CodeExercises, of roughly 180M tokens of Python exercises and solutions, also generated by GPT-3.5.
The filtering step is where the human judgment lives. They used GPT-4 to annotate about 100k samples for educational value, then trained a random forest classifier on those labels and ran it over the rest of the corpus. So GPT-4 sees a small slice, a cheap classifier does the heavy lifting, and GPT-3.5 writes the synthetic portion. That division of labour is the part we expect other groups to copy first, because it is cheap.
Where the jump actually happens
The number that surprised us is the base model. phi-1-base, trained only on the CodeTextbook mix without any finetuning, gets 29% on HumanEval. That is already respectable for 1.3B parameters. The finetune on 180M tokens of exercises takes it from 29% to 50.6%. Most of the headline gain comes from a dataset smaller than many people's evaluation suites.
The authors say the finetuning unlocks capabilities that the exercises do not obviously contain, including better use of external libraries that do not appear in CodeExercises. We believe the observation and we are unsure of the mechanism. One reading is that the exercises teach the model to produce a complete function body in the format HumanEval expects, which is a format effect rather than a capability effect. The paper does not fully separate those.
They also trained a 350M version, phi-1-small, on the same data. It gets 45% on HumanEval. That gap between 350M and 1.3B is smaller than we would have guessed, which supports their thesis that the data is doing the work, and it also suggests they may be near the ceiling of what this particular corpus can teach.
The contamination question
Any time a synthetic dataset generated by a strong model produces a big jump on a public benchmark, the first thing to check is whether the benchmark leaked into the data. The authors know this and spent a section on it. Their n-gram analysis found only four cases of 13-gram overlap with HumanEval, all of which they judge to be false positives.
N-gram overlap is a weak test for a model-generated dataset, since GPT-3.5 can paraphrase a HumanEval problem without reusing its exact words. The better experiment in the paper is the pruning one. They removed problems from CodeExercises that were semantically similar to HumanEval problems, cutting between 42.5K and 354K entries out of 879.5K depending on the threshold, and retrained. Even after removing more than 40% of the exercises, the retrained phi-1 still outperformed StarCoder.
That is a reasonable defence and it is not airtight. The pruned model's absolute HumanEval score drops, and the comparison the paper leans on is against StarCoder rather than against the unpruned phi-1. We would want to see the same pruning done against MBPP, and we would want a held-out set of fresh problems written after the training data was frozen. Those are things an independent group could do with the released model.
What the paper does not claim
The paper is narrower than the title. It is about Python, and about short function-level completion. The authors list sensitivity to prompt wording as a limitation and note that performance drops significantly as prompts get longer. A 1.3B model trained on 7B tokens has not seen enough of the world to be a general coding assistant, and they do not pretend otherwise.
It is also a paper about one generator. Everything synthetic here came from GPT-3.5, so the ceiling of phi-1 on any task is bounded by what GPT-3.5 can produce reliably. If you view this as distillation with extra steps, the interesting question becomes how much of the teacher's ability survives when the student is 100 times smaller. Half of HumanEval surviving is a striking answer.
Why this matters for the scale argument
The scaling laws we have all been quoting treat tokens as interchangeable. This paper is a direct challenge to that assumption. If a carefully filtered and partly synthetic 7B token corpus can substitute for a much larger raw one on a narrow task, then the constant in front of the data term in those laws is not a constant. It depends on how the data was made.
The question we care about is whether this generalizes beyond code. Code has a verifier, in the sense that you can run it, and exercises with solutions are easy to generate and check. Natural language reasoning has no equivalent oracle. Our guess is that the same group will try the recipe on general text next, and that the result will be less clean, because the filtering signal will be weaker.
What we would want someone to try is a replication with a different generator and a different filtering model, to see how much of the result is specific to GPT-3.5 and GPT-4. If a group using open models can get a 1.3B student above 40% on HumanEval with a fully documented pipeline, the small model, good data bet is real. Until then it is one impressive result from one lab with one teacher.
Sources
From the foundation