Can a state space model do recursive reasoning? A 7M parameter test
Wang and Reid swapped Mamba-2 blocks into the Tiny Recursive Model and matched or beat the attention version on ARC-AGI-1 at equal size. A replication note on the tiny recursive lineage and on why the loop seems to matter more than what sits inside it.
The lineage in three numbers
The Hierarchical Reasoning Model showed last year that 27 million parameters could reach 40.3 percent on ARC-AGI-1 by refining a latent state in a loop instead of producing a chain of thought. Alexia Jolicoeur-Martineau's Tiny Recursive Model then cut that to a single two-layer network of about 7 million parameters and improved the score to roughly 45 percent on ARC-AGI-1 and 8 percent on ARC-AGI-2, trained on around a thousand examples. Those two results established that the recursion was doing the work, since removing a whole second network made things better.
Wenlong Wang and Fergal Reid at Intercom have now asked the next question. If the loop is the important part, does the block inside the loop have to be a Transformer? Mamba-2's state update is itself a recurrence, so it is a natural candidate. They kept the TRM scaffold exactly, three outer cycles and four to six inner cycles, the same two latent states and the same output heads, and replaced the attention blocks with a Mamba-2, Mamba-2, attention, MLP pipeline. Parameters were matched at 6.86 million against 6.83 million for the original.
What changed on ARC-AGI-1
On pass at 2, which is the official ARC metric, the hybrid scores 45.88 percent against 43.88 for the attention baseline, a gain of two points. At pass at 1 the two are level, 40.50 against 40.75. The gap widens as K grows, reaching 4.75 points at pass at 100 and 4.25 at pass at 1000. The training curves show the same pattern from early in the run, so this is a consistent property of the architecture rather than a lucky final checkpoint.
The authors read the pass at K shape as a coverage versus selection trade-off, and their diagnostic statistics support it. Under the ARC protocol each test input is expanded into about 880 augmentations, each produces a prediction, and the answers are voted. The hybrid generates 339.5 unique candidates per puzzle against 266.6 for the baseline and has higher vote entropy, 5.39 bits against 4.56. The baseline concentrates 41.1 percent of votes on its top candidate against 32.9 for the hybrid. So the Mamba version explores more and commits less, which helps whenever the correct answer is rare in the pool and hurts slightly when it is already dominant. Stratified by difficulty, the hybrid gains 4.9 points at pass at 5 on hard inputs and the baseline gains 4.6 points at pass at 1 on easy ones.
The other two tasks are less tidy
On Sudoku-Extreme the picture inverts. Both attention-style models lose to the MLP-t variants, which mix across positions with a dense transposed MLP. TRM-mlp-t gets 87.4 percent exact accuracy, the Mamba plus MLP-t hybrid 84.2, the original attention TRM 72.2, and the Mamba plus attention hybrid 66.5. The authors' reading is that a fixed 9 by 9 grid rewards all-to-all dense mixing over either selective attention or sequential state.
On Maze-30x30-Hard both MLP-t variants score exactly zero, and the Mamba plus attention hybrid reaches 80.6 percent against 60.8 for the attention baseline. The paper is candid that training on this task fluctuated between 6 and 85 percent across checkpoints, so the Maze numbers are preliminary. What survives across the three tasks is the narrow claim. Putting Mamba-2 into the recursive scaffold does not break reasoning, and on the benchmark that matters most it improves candidate coverage.
A note on what a replication in this lineage costs
These models are tiny and the training runs are not. Ronan McGovern's test-time adaptation paper from November reports that reproducing the TRM ARC-AGI-2 run takes roughly 48 hours on four H100s for 700,000 or more optimiser steps, which is far outside the ARC Prize competition budget of four L4s for twelve hours. His replication reached 10 percent on the public ARC-AGI-2 evaluation set, above the 7.8 percent in the original paper, and he attributes the difference to run-to-run variance. Wang and Reid also chose not to reproduce the TRM-mlp-t result on ARC to save compute, which is a reasonable choice and also a reminder that a single comparison at 6.8 million parameters still costs days of GPU time.
The variance point deserves weight. When a lineage's headline numbers move by two percentage points between papers, and the same code moves by two percentage points between seeds, a two point gain needs several seeds to be believed. The pass at K curve and the candidate statistics are more convincing to us than the pass at 2 headline precisely because they are consistent along a whole axis rather than a single cell.
Why the loop seems to matter more than the block
Three architectures have now gone through the same recursive scaffold, a two-network hierarchy, a single Transformer and a Mamba-2 hybrid, and all land within a few points of each other on ARC-AGI-1 at under 30 million parameters. Meanwhile the choice of mixing operator flips the ranking completely between Sudoku and Maze. The stable ingredient across all of it is the outer loop with post-norm on each residual add, which the authors argue is what keeps the hidden state bounded across unrolled iterations regardless of what the operator is.
The experiment we would like to see is the one the paper gestures at in its discussion. Mamba-2 has its own recurrence inside each block. If the outer loop is what matters, it should be possible to fold some of that iteration into the state update and reduce the number of outer cycles without losing accuracy. If that works, we would learn something about whether latent recursion is a property of the scaffold or of the computation, and it would be a clean test at a scale where anyone with four GPUs and a weekend can try.
Sources
From the foundation