A sharp line in the trace

When a reasoning model works through a maths problem it produces a long trace, and the assumption behind test-time compute is that the trace is where the work happens. A paper posted on June 12 by Daniel Scalena, Sara Candussio, Luca Bortolussi, Elisabetta Fersini, Malvina Nissim and Gabriele Sarti asks which parts of the trace actually determine the answer. Their finding is that there is a commitment boundary, a sharp transition from transient intermediate guesses to a stable high-confidence answer, and that it typically happens in a single step. Everything after that step is, in their term, epiphenomenal. It reads like reasoning and it does not move the answer.

The method is direct. Truncate the trace at successive sentence boundaries, force the model to answer with greedy decoding from that point, and watch when the answer stops changing. That gives you a per-problem location for the boundary. The authors then check that the later steps are causally inert rather than merely redundant, which is the difference between text that could have been cut and text that the model was ignoring.

How much of the chain is after the line

The study covers three open model families, gpt-oss-20b, Qwen3-14B and gemma-4-26B-A4B-it, on four benchmarks, MATH-500, AIME 2025, ZebraLogic and GPQA-Diamond. The amount of post-commitment text varies a lot by model. Gemma converges after 13 to 23 percent of its tokens depending on the dataset, so roughly four fifths of what it writes comes after it has already decided. The gpt-oss model needs 43 to 68 percent of its tokens before the answer settles. Across models and datasets the headline is that chains could be shortened by up to 55 percent on average with negligible impact on performance.

The variation matters as much as the average. A model that commits at 15 percent and keeps writing is spending most of its inference budget on text with no causal role. A model that commits at 60 percent is using more of its budget, which is either more honest reasoning or less efficient reasoning, and the paper's method cannot by itself tell you which. What it can tell you is where the line falls for a given model on a given task, and that is enough to act on.

Detecting the boundary without truncating

Truncating at every sentence is expensive, so the second contribution is a way to detect the boundary from activations. The authors train lightweight attention probes on model internals to predict whether the answer has stabilised. The probes reach above 90 percent detection on held-out data and hold up across in-distribution and out-of-distribution reasoning tasks. That turns a diagnostic into a mechanism. You can run the probe during generation and stop when it fires.

The early-exit numbers are the practical result. With a tolerance window of one step, meaning the exit is allowed to be one sentence off, the method recovers 89 percent of full-chain accuracy while saving 39 percent of chain-of-thought tokens. On out-of-distribution datasets the savings range from 6 to 35 percent with small accuracy losses, and the worst case is an 11 percent decline on ZebraLogic. The paper summarises the probe-mediated approach as saving up to 35 percent of tokens with minimal degradation. So the 55 percent figure is the ceiling from oracle truncation, and 35 to 39 percent is what a deployable probe gets you today.

What this does to monitoring

The result we find most uncomfortable is for chain-of-thought monitoring. The case for reading traces as a safety signal rests on the trace reflecting the computation that produced the action. If a large fraction of the trace is written after commitment, then the parts of it that look like checking, reconsidering, or weighing alternatives may be narration of a decision already made. A monitor that sees the model write "let us verify this" after the boundary is watching text with no causal connection to the output.

A worked case. Suppose a model on a coding task commits to a shortcut at sentence twelve of a forty-sentence trace, and sentences thirteen through forty discuss why the shortcut is safe. A monitor reading the full trace sees a considered decision. The paper says the considering happened in sentence twelve or before, and what follows is post hoc. Where the model commits is the only place the reasoning can be read as reasoning. That does not make the trace useless. It means the monitor needs to know where the boundary is, and this paper gives a way to find it.

What it does to the cost story

Test-time compute is sold as buying accuracy with tokens. This paper says a large and model-dependent share of those tokens buy nothing. If you can cut 35 percent of reasoning tokens with a probe and lose little, then the price per task of a reasoning model is inflated by about that much, and the provider has an incentive to leave it inflated when they bill by output token. We would expect labs to adopt something like the probe internally whether or not they publish it, and we would expect the effect to show up as quieter models rather than cheaper ones.

The experiment we want is the reverse. Take the post-boundary text and change it. Insert a wrong verification step after commitment and see if the answer follows. If it never does, the epiphenomenal label is exact and monitors should be told to weight late text near zero. If it sometimes does, the boundary is soft in a way the truncation experiments would miss, and we need to know how often. Either result would be useful, and both are cheap to run with the authors' setup.

Sources

  1. Scalena et al., Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models (arXiv 2606.13603)