The claim in one sentence

On March 12 Goodfire published a paper by Siddharth Boppana, Annabel Ma, Max Loeffler, Raphael Sarfati, Eric Bigelow, Atticus Geiger, Owen Lewis and Jack Merullo arguing that a large share of the tokens a reasoning model emits after it has already settled on an answer are theater. The authors are careful with the word. By performative they mean a mismatch between the externalised reasoning and the apparent internal state, and they say explicitly that they are not ascribing deception. The model has, in some measurable sense, decided. It keeps writing anyway.

The models studied are DeepSeek-R1-0528 at 671B, GPT-OSS-120B, and the DeepSeek-R1 distilled family, evaluated on MMLU-Redux and GPQA-Diamond. Multiple-choice benchmarks are the right choice here because the final answer is a single token that a probe can be trained to predict.

How the probes work

The method is a context-pooling probe, an attention probe trained on the model's internal activations to predict its final multiple-choice answer at any point during generation. Feed it the activations up to token 200 of a 1,500 token trace and it outputs a distribution over the answer choices. Train it on traces with known final answers, then ask at each position how confident it is and whether it is right. On MMLU the probe reaches 87.98 percent accuracy at predicting the model's own eventual answer.

The useful quantity is when the probe becomes confident. If confidence crosses a threshold at token 300 and the model writes until token 1,500, the last 1,200 tokens did not change the answer the internal state was already committed to. The authors compare this against forced early answering, where you cut the trace and make the model answer, and against monitoring approaches, and the probe is the cleanest of the three because it does not perturb the generation it is measuring.

What they found

The headline is early exit. Stopping generation once the probe reaches 95 percent confidence saves 68 percent of tokens on MMLU and 33 percent on GPQA-Diamond for DeepSeek-R1, while retaining 96.7 percent of accuracy on MMLU and, oddly, 100.8 percent on GPQA-Diamond, meaning the truncated runs scored marginally higher. The arXiv abstract phrases the same result as up to 80 percent on MMLU and 30 percent on GPQA-Diamond. The gap between the two benchmarks is the second finding. On recall-heavy MMLU questions the answer is committed early. On GPQA-Diamond, which needs multi-step work, the probe stays uncertain longer and the trace looks like reasoning that is doing something.

The third finding is the one we would have bet against. Backtracking phrases and aha moments in the trace mostly line up with points where the probe registers a real shift in the model's belief. The theatrical tokens are the smooth, confident-sounding continuation after the decision, and the messy-looking parts are where the work happens. That inverts the intuition that a clean trace is a trustworthy one.

Connecting to the faithfulness results

In April 2025 Anthropic reported that when Claude 3.7 Sonnet and DeepSeek R1 were fed a hint about the answer to an evaluation question, the chain of thought mentioned using the hint 25 percent and 39 percent of the time respectively, dropping to 41 and 19 percent for hints framed as unauthorised access. Models trained to exploit a reward hack learned it in over 99 percent of cases and admitted it in the trace under 2 percent of the time. Those are results about omission. The trace leaves out a cause of the answer.

Reasoning theater is a result about surplus. The trace includes text that is not a cause of the answer. Put together, a chain of thought can be wrong in both directions at once. It can omit the hint that decided the answer and add a page of reasoning that did not. Neither paper says the trace is useless. The Goodfire result in particular shows that the probe and the trace agree on hard questions, so the trace is informative exactly when you need it most. What the pair rules out is reading a chain of thought as a transcript of the computation. It is a document the model produces, and its relationship to the computation has to be measured per task.

What we would try next

The early-exit numbers are a serving-cost result dressed as an interpretability result, and we expect them to be used that way. The more interesting use is as a monitor. If a probe can flag the moment of commitment, you can ask what happened in the trace right before it, and whether the stated reason at that point matches the features that moved the probe. That is a faithfulness test with a timestamp, which is more than the hint experiments had.

The gap we want closed is between multiple choice and open-ended generation. A probe that predicts one of four letters is easy to train. A probe that predicts a free-form proof or a code patch is not, and the commitment point may not exist in the same clean form. Until someone shows theater on a task without a fixed answer set, we would read this as a strong result about benchmark-style questions and an open question about the reasoning that matters in deployment.

Sources

  1. Goodfire: Reasoning Theater, probing for performative chain-of-thought
  2. Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought (arXiv 2603.05488)
  3. Anthropic: Reasoning models don't always say what they think