A 27 million parameter model and the ARC-AGI headline
Sapient's Hierarchical Reasoning Model reported 40 percent on ARC-AGI-1 with 27M parameters and about a thousand training examples. The ARC Prize team ran it on the hidden set and ablated it. The hierarchy did little. The refinement loop and the augmentation did most of the work.
The claim
The Hierarchical Reasoning Model paper from Guan Wang and colleagues describes a recurrent architecture with two modules, a high level module that updates once per cycle and is meant to do slow abstract planning, and a low level module that updates every timestep and does the detailed computation. The whole thing is about 27 million parameters, both modules are encoder-only transformer blocks, and it runs in a single forward pass with no chain of thought. It is trained on about 1,000 examples per task with no pretraining.
The headline numbers are what made it travel. The paper reports 40.3 percent on ARC-AGI-1 against 34.5 percent for o3-mini-high and 21.2 percent for Claude 3.7 with an 8K context. It also reports near perfect accuracy on Sudoku-Extreme and on finding optimal paths in 30 by 30 mazes, tasks where chain of thought methods scored zero. The framing leans on the human brain, with hierarchical and multi timescale processing cited as the inspiration. A tiny model beating frontier reasoning systems on a benchmark designed to resist memorisation is a very good story, and it spread accordingly.
What the ARC Prize team measured
ARC Prize has a semi-private evaluation set that models cannot have seen, and they ran HRM on it. The verified score on ARC-AGI-1 was 32 percent, down from the 41 percent the team reproduced on the public set. On ARC-AGI-2 it scored 2 percent. The runs took 9 hours 16 minutes and 12 hours 35 minutes respectively, at about 1.48 and 1.68 dollars per task. So the result held up in the sense that a 27M model really does get a third of ARC-AGI-1 right on unseen tasks, which is remarkable for the size. It did not hold up on the harder benchmark at all.
The more interesting part is the ablation. The team swapped the hierarchical two module design for a regular transformer of similar size, with no hyperparameter tuning, and came within about 5 percentage points of HRM. That is the brain inspired component, and it is worth roughly 5 points. The outer refinement loop, where the model iterates on its own prediction, was worth far more. Going from no refinement to a single refinement step gave a 13 point jump, and performance kept climbing up to 8 loops. They also found that training with refinement mattered more than refining at inference time.
Augmentation and the puzzle_id embedding
The paper trains with heavy augmentation of the ARC grids, translations, rotations, flips and colour permutations. ARC Prize found that only 300 augmentations per task got near maximum performance, and 30 augmentations, 3 percent of what the paper used, came within 4 percent of the maximum. So the augmentation is doing real work, but most of it is done early and the rest of the budget is mostly spent for nothing.
The finding that changes how you should read the headline is about what the model is actually trained on. HRM learns a puzzle_id embedding for each task, and it is trained on the evaluation tasks' demonstration pairs at evaluation time. ARC Prize trained a version only on the 400 evaluation tasks and got 31 percent, versus 41 percent for the full setup. In their words, most of the performance is driven by training on the tasks seen at evaluation time. This is transductive prediction, fitting the specific puzzle in front of you, rather than inductive program synthesis that would transfer to a new puzzle. That is allowed under ARC rules, since the demonstration pairs are part of the task, but it is a different thing from what the paper's framing implies.
What survives
Plenty survives, and we want to be fair to the paper. A 27M parameter model with no pretraining scoring 32 percent on the hidden ARC-AGI-1 set is a real data point about how much of that benchmark is reachable with test time training and iteration rather than scale. The one step gradient approximation, deep supervision with detached hidden states between segments, and the adaptive computation scheme that learns when to halt are all engineering that a replicator can reuse. The refinement loop result in particular is the kind of finding that should be pulled out and studied on its own.
What does not survive is the causal story. The paper attributes the result to the hierarchical, brain inspired structure, and the ablation says that structure is worth about 5 points out of 40. The brain framing is what got the paper shared, and it is the part with the least evidence behind it. This is not unusual. Architectural papers often carry an explanatory narrative that outruns the ablations, and a benchmark result gets read as a confirmation of the narrative when it is only a confirmation of the number.
The general lesson for reading ARC results
When a small model does well on ARC, the first question should be whether it is trained per task or once. The second should be whether the score came from a refinement loop, from augmentation, or from the architecture the paper is named after. ARC Prize did those checks and it took them an ablation study, a hidden set run and a modest amount of compute. That is cheap relative to the attention the paper received, and it is the kind of work we would like to see done routinely rather than only when a result goes viral.
The experiment we would run next is the plain transformer plus outer refinement loop, trained transductively, with 30 augmentations. If that matches HRM within noise, the field has a simpler and better understood baseline for this class of result, and any future architecture claim on ARC has something honest to beat.
Sources
From the foundation