GSM-Symbolic: the reasoning paper everyone cited for the wrong reason
Apple researchers turned 100 GSM8K problems into templates, regenerated them with new names and numbers, and watched accuracy fall. The paper was read as proof that models cannot reason. What it measures is narrower and more useful.
What they built
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio and Mehrdad Farajtabar at Apple took 100 questions from the GSM8K test set and rewrote each as a template with variables for the names and the numbers. Sampling each template 50 times gives 5,000 fresh problems that have the same logical structure as the originals and none of the same surface text. They call the result GSM-Symbolic and evaluate 25 models on it, from 2B open models up to GPT-4o and o1-preview.
The paper posted on 7 October, and within days it was being cited as evidence that language models do not reason. The authors' own conclusion is close to that: they say current models attempt to replicate reasoning steps from training data rather than performing logical reasoning. We want to separate what the experiments show from that interpretation, because the experiments are good and the interpretation is doing more work than they can bear.
Result one: GSM8K scores sit at the top of the distribution
The first finding is about the benchmark rather than the models. For 21 of the 25 models, the original GSM8K accuracy falls at the right edge of the distribution of accuracies across the 50 GSM-Symbolic instantiations. Gemma2-9B goes from 87 percent on GSM8K to 79 percent on average across the symbolic variants. Phi-3-mini drops from 85 to 81. GPT-4o barely moves, at 95 in both.
If a model's score on the specific published problems is reliably better than its score on structurally identical problems with different numbers, the published problems are easier for it in a way that has nothing to do with their structure. The simplest explanation is contamination, direct or through the many paraphrases of GSM8K on the web. That is a fact about GSM8K as a measuring instrument, and it is the most useful thing in the paper.
Result two: numbers hurt more than names
The second experiment separates the two kinds of change. Swapping only the names in a problem produces a modest spread of accuracies. Swapping only the numbers produces a much larger one. On a pure logical view of the task these should be equivalent, since neither changes the reasoning required.
This is the result that got read as models cannot reason. We read it differently. The same models get the vast majority of these problems right under every instantiation. What varies is a margin of several points that depends on which particular numbers appear. That is sensitivity to surface form, and it is a real weakness, but a system that is right 79 percent of the time regardless of the numbers is doing something that a pure retrieval of memorised steps would not do. The paper's own difficulty ladder shows the same thing: dropping a clause (GSM-M1) raises accuracy and adding one or two (P1, P2) lowers it and widens the variance, with Phi-3-mini going from 88 on M1 to 45 on P2 and Gemma2-9B from 87 to 42. Accuracy that degrades smoothly with the number of reasoning steps is what you would expect from a fallible reasoner as well as from a pattern matcher.
Result three: GSM-NoOp
The third experiment adds one sentence to each problem that looks relevant and is not, such as noting that some of the kiwis were smaller than average in a problem about counting kiwis. On this GSM-NoOp set, accuracy collapses. Phi-3-mini drops by 65 points. Llama3-8B goes from 75 to 19. o1-preview goes from 77.4 to 19. Models convert the irrelevant statement into an arithmetic operation, subtracting the small kiwis, and eight in-context examples of the same question with the correct answer do not fix it.
This is the strongest result in the paper, and what it exposes is a specific learned prior: every number in a grade-school word problem is there to be used. That prior is correct on GSM8K, where problems are written so that every number matters, and models trained on a lot of GSM8K-like data have absorbed it. A test that violates the convention catches them. Whether a model that had seen distractor-laden problems in training would fall the same way is exactly the question the paper does not answer.
What the paper is actually good for
Read as a paper about benchmarks, GSM-Symbolic is excellent. It gives a cheap procedure for turning any templated benchmark into a distribution of scores rather than a point, and the width of that distribution is a number every leaderboard should report. It shows that the published instance of a benchmark is systematically flattering. And it shows that an adversarial clause can expose a shortcut that clean instances hide.
Read as a paper about whether models reason, it proves less than its abstract suggests, because the same evidence is consistent with a reasoner that has a bad prior about distractors and a sensitivity to numeric surface form. The experiment we would run next is to train on distractor-laden templates and test on held-out ones. If NoOp accuracy recovers, the collapse was a data artefact. If it does not, the authors' stronger claim gets a lot more credible, and we would want to know that.
Sources
From the foundation