GSM1k: rebuilding a benchmark to find out who overfit it
Scale AI commissioned a fresh set of grade-school maths problems matched to GSM8K and found that some model families drop up to 13 points on it. A replication story about what a held-out twin dataset can reveal about contamination, and what it cannot.
The idea
GSM8K has been the default test of grade-school reasoning for three years, and by now its problems and their solutions are on the open web in thousands of places, including fine-tuning sets built from it. So when a model scores well on GSM8K, there are two explanations, and the leaderboard cannot tell them apart. Scale AI's paper, posted on May 1 by Hugh Zhang and fourteen coauthors, attacks the problem the direct way. They wrote a new benchmark that is meant to be statistically indistinguishable from GSM8K and then compared every model's score on the old set to its score on the new one.
The logic is simple enough that the whole paper stands or falls on execution. If a model has genuinely learned to do this kind of arithmetic, it should score about the same on two draws from the same distribution. If it has memorised the first draw, it will score higher there than on the second. The gap between the two scores is the signal.
How you build a twin
The execution is most of the paper and most of the cost. Human annotators wrote the problems without any language model assistance, and the authors describe roughly 2,108 candidate problems entering the pipeline, 1,419 surviving independent verification by a second annotator who solved them blind, and 1,375 passing a final quality audit, from which the released evaluation set of 1,205 problems was drawn. Annotators were paid 20 to 25 dollars an hour and the whole exercise cost about 180,000 dollars.
Matching was done on the things that make a maths problem hard. The distribution of the number of arithmetic operations per solution and the magnitude of the answers were matched to GSM8K. Then the authors ran two checks that we think are the strongest part of the paper. Annotators given problems from both sets could tell which set a problem came from only 21.83 percent of the time, which is close to chance. And in timed 15-minute sessions, annotators solved 4.07 GSM8K problems on average against 4.36 GSM1k problems, with overlapping spreads, so the new set is, if anything, slightly easier for humans. That rules out the boring explanation for any gap, which is that the new problems are simply harder.
Who dropped and who did not
The headline number in the first version of the paper is a drop of up to 13 percent from GSM8K to GSM1k across the models tested, and the pattern is what makes it interesting. Two families, Phi and Mistral, showed evidence of systematic overfitting across almost all model sizes. In the detailed results, Phi-2 falls 6.3 points, Math-Shepherd-Mistral-7B-RL falls 7.2, and Yi-6B-Chat has the largest gap at 8.0. Smaller effects show up in Xwin, Gemma and CodeLlama.
At the other end, the frontier models barely moved. GPT-4, Claude and Gemini showed minimal signs of overfitting, and Mistral Large was the only member of the Mistral family with no gap at all. The authors read this as a sign that stronger reasoning lets a model generalise even when it has seen the test data, and we think that is one plausible reading. Another is that the frontier labs are more careful about decontamination, or that GSM8K is a vanishingly small fraction of a frontier pretraining set. The paper cannot separate those, and it does not claim to.
The paper does offer one piece of mechanism. Across models, the probability a model assigns to generating an actual GSM8K example correlates with its performance gap, with a Spearman r squared of 0.32 in the first version. That is a moderate relationship. It supports the story that some of the gap is literal memorisation of the test set, and it leaves room for gaps that come from something else, such as training on paraphrases or on the same templates without the exact strings.
What a twin dataset cannot tell you
The method has a ceiling and the authors are clear about it. A held-out twin can show that a model's GSM8K score overstates its ability on GSM8K-like problems. It cannot show whether that ability transfers to anything else, because the twin is by construction a copy of the original distribution. A model that is not overfit to GSM8K may still be overfit to the genre of short word problems with integer answers, and GSM1k would never see it. Every model in the study, including the ones with large gaps, still solved most of the new problems, and the authors point out that this shows real generalisation to problems guaranteed absent from training.
There is also a durability problem. The paper's whole value depends on GSM1k staying out of training data, so Scale is not releasing it. The stated conditions for release are that three open-source models reach 95 percent accuracy on it, or June 2025, whichever comes first. The evaluation code will be open. That is the right choice for the measurement and the wrong one for reproducibility, and we do not see a way to have both. A benchmark that nobody can see is a benchmark you have to take on trust, and a benchmark everyone can see is one you can only trust once.
What we would like to see repeated
The obvious next step is to run the same procedure on the other benchmarks that everyone fine-tunes toward. A MATH twin would be more expensive because the problems are harder to write and verify, but the same drop-and-compare design would work. A HumanEval twin would be cheaper and probably more revealing, given how many code fine-tuning sets have been built from it. The procedure is the contribution here. The specific numbers will be stale within a year.
The result we would most like to see is a longitudinal one. If GSM1k is held back until mid-2025 and then released, someone should re-run every model from this paper on it on the day it goes public and again a year later, once it has had time to leak. That would give the first direct measurement of how fast a benchmark rots, which is the number every evaluation paper needs and none of them have.
Sources
From the foundation