The Illusion of Thinking and the illusion of the rebuttal
Apple's puzzle paper reported reasoning models collapsing past a complexity threshold. Within days a rebuttal argued the collapse was output limits and unsolvable River Crossing instances. Both sides show how easily an evaluation artifact passes for a capability finding.
What the paper measured
Shojaee, Mirzadeh, Alizadeh, Horton, Bengio and Farajtabar at Apple set out to avoid contaminated math and coding benchmarks by evaluating reasoning models on puzzle environments where complexity can be dialled up one notch at a time. Tower of Hanoi from one disk to twenty is the clearest example. They compared reasoning models against non-reasoning versions of the same family, DeepSeek-R1 against DeepSeek-V3 for instance, under matched compute budgets.
Three regimes came out. On easy puzzles the non-reasoning models matched or beat the reasoning ones. On medium puzzles the reasoning models pulled ahead. Past a threshold, both collapsed to zero. The paper reports a second, stranger finding: reasoning effort, measured in thinking tokens, rose with complexity up to a point and then fell, even though the models had token budget left. The authors read this as a scaling limit inherent to the approach. Handing the models the solution algorithm did not help much either.
The rebuttal
On June 10, a short comment by A. Lawsen appeared on arXiv, revised on June 16. It makes three claims. First, the Tower of Hanoi experiments at high disk counts require move lists longer than the models' output token limits, and models say so in their outputs before the grader marks them wrong. Second, the automated scoring cannot tell a reasoning failure from a model that stopped because it ran out of room. Third, the River Crossing environment includes instances with more than five actors that are mathematically unsolvable given the boat capacity, and models were scored as failing them.
The comment also reports that when models were asked to produce a generating function for the Hanoi moves rather than the exhaustive list, preliminary tests across several models showed high accuracy on instances the paper had recorded as complete failures. On that basis it argues the collapse is an experimental artifact.
The rebuttal has its own artifacts
We read the comment twice and was persuaded on the River Crossing point, which is a plain error in the benchmark. Unsolvable instances scored as failures is not a matter of interpretation. The token-limit argument is weaker than it looks. It is true for the largest Hanoi instances, and it does not explain the drop in reasoning effort before the limit is reached, which the Apple paper documents and the comment does not engage with. A model that gives up with budget remaining is doing something that a token cap cannot explain.
The generating-function result is the part that worries us most, and it is the part that spread fastest. Asking for a function instead of a move list changes the task from executing an algorithm to recalling one. Tower of Hanoi's recursive solution is in every introductory programming course and therefore in every training corpus. High accuracy at writing it down tells you the model remembers it. The original paper had already shown that supplying the algorithm did not rescue execution. So the rebuttal answers a question the paper did not ask, and the answer is then reported as overturning the paper.
The refusal reading
Sean Goedecke posted a critique on June 8 that we find more useful than either paper. He ran Apple's prompts against DeepSeek-R1 and quoted the start of a trace on the ten-disk puzzle. The model computes that 1023 moves are needed, states that generating them manually is impossible, and starts hunting for a shortcut. His reading is that past eight or nine disks the skill being tested silently changes from working through the sequence to inventing a way around it. The models are not trying and failing. They are declining to start.
That reading explains the falling token counts without invoking either a scaling limit or an output cap. It also suggests a cheap experiment nobody in this exchange ran: instruct the model, forcefully, to enumerate every move, and see whether the collapse moves. Goedecke notes the model grumbles about tedium even at low disk counts, so persistence is a trainable property and probably not a fixed one. He also concedes, in an edit after the Reddit discussion, that the other puzzles failed earlier without the same explosion in step count, which his account does not cover.
Artifacts in both directions
Here is what we think happened. Apple built a clean, scalable benchmark and did not check hard enough whether its hardest instances were solvable or expressible within the output window. The result got read as reasoning models cannot reason. Then a rebuttal identified real flaws, added a task substitution that measures recall rather than execution, and got read as the paper was wrong. Each side produced an evaluation artifact and each was received as a capability finding.
The correction is boring and mechanical. Before claiming a threshold, verify that every instance past it is solvable and that a correct answer fits in the output. Before claiming a rescue, verify that the new prompt measures the same skill. Report the models' own statements about giving up as data rather than noise. We would like to see the Hanoi experiment rerun with three conditions, exhaustive list with a forced-persistence prompt, generating function, and code that produces the list, on the same instances, with the failure reason for each run recorded by hand. Until someone does that, the honest summary is that reasoning models stop executing long algorithms somewhere around a thousand steps, and we do not yet know whether that is a limit or a preference.
Sources
From the foundation