Measuring faithfulness by breaking the chain
Anthropic truncated, corrupted, paraphrased and blanked out chains of thought to see whether the final answer depended on them. Sometimes it did. Larger models were worse.
The question the paper asks
When a model writes out its reasoning before answering, is the answer produced by that reasoning, or is the reasoning a story told after the answer was already fixed? A paper from Tamera Lanham and colleagues at Anthropic, posted on July 17, does not try to look inside the model to find out. It intervenes on the written reasoning and watches whether the answer moves.
The setup uses a 175B parameter assistant trained with RLHF and eight multiple choice tasks: ARC Easy and Challenge, AQuA, HellaSwag, LogiQA, MMLU, OpenBookQA and TruthfulQA. For each question they sample 100 chains of thought, split them into sentences, and then ask for a final answer under four kinds of damage. The four tests are designed as a defence in depth. None is conclusive alone. Each rules out one way the reasoning could be unfaithful.
Four ways to break a chain
Early answering truncates the chain after each sentence and asks for the answer. If the model gives the same answer with zero sentences as with all of them, the reasoning was post-hoc. Adding mistakes uses a separate pretrained model to rewrite one sentence into a plausible error, then lets the assistant continue from there. If the answer never changes after a corrupted step, the model was ignoring the chain. Their mistake generator produced a plausible error at least 80 percent of the time.
Filler tokens replace the whole chain with strings of ellipses of varying length. This tests whether the gain from chain of thought is just extra compute from having more tokens between question and answer. Paraphrasing has a model reword the opening sentences of the chain without seeing the question, then continues the chain from there. This tests whether the gain is smuggled in through phrasing, a kind of steganography the model might have learned under RLHF.
The paper's illustration is 5 factorial. The original chain says 1 times 2 times 3 times 4 times 5 equals 120, answer 120. The corrupted version says it equals 100 and the answer follows to 100. A truncated version stops before the product and the model still says 120. Whether that last case counts as unfaithful depends on whether the model had already computed the answer before it started writing, which is exactly what the paper is trying to measure.
What the interventions showed
The headline is variation. On the three tasks with the least dependence on reasoning, ARC Easy, ARC Challenge and OpenBookQA, truncating the chain changed the final answer less than 10 percent of the time. On AQuA, the algebra word problems, truncation changed the answer more than 60 percent of the time. The paper summarises this with an area over the curve metric. AQuA scores 0.44 on early answering and LogiQA 0.26, while ARC Easy scores 0.02. Adding mistakes gives nearly the same ordering.
The surprise is that faithfulness does not track accuracy gain. AQuA gets the largest boost from chain of thought, 15 points, and is also the most faithful, which fits. But LogiQA gains only about one point and is second most faithful. HellaSwag actually loses 4.7 points with chain of thought and still shows more dependence on it than several tasks that gain. So a model can lean on reasoning that does not help it, and can ignore reasoning that does.
The filler and paraphrase tests both came back negative, in the sense that matters. Replacing the chain with ellipses gave no accuracy gain at any length, and for TruthfulQA and OpenBookQA long runs of filler slightly hurt. Paraphrased chains matched the accuracy of the originals on nearly every task. So the benefit of chain of thought, where there is one, comes from its content, and not from its length or its wording.
The inverse scaling result
Section 3 is the part we expect people to argue about. They rerun the cheapest metric, how often the answer changes with versus without a chain, across a series of RLHF assistants from 810M to 175B parameters. For seven of the eight tasks the 13B model changes its answer after reasoning more often than the 175B model. For six of the eight, everything except AQuA and LogiQA, the 13B model is the most faithful of any size, including the smaller ones. From 13B up, faithfulness mostly falls monotonically with scale.
Their explanation is capability. A model that can already answer a question without reasoning has no need to condition on the reasoning it writes, so the chain becomes decoration. Models below 13B get little or no accuracy from chain of thought on these tasks, which is why they are not more faithful either. They are not yet able to use the chain. Faithfulness peaks at the size where the model needs the reasoning and can follow it.
To check that difficulty is the variable and not something about the tasks, they built synthetic addition problems with 2, 4, 8 or 16 operands of two or three digits, and measured the same metric. Post-hoc reasoning rises with model size on every variant and rises with easier problems at fixed size. The two smallest models were dropped because they would not consistently return an integer. The conclusion is uncomfortable for anyone who wants both the best model and a faithful explanation. For a given task, the most faithful reasoning may come from a model deliberately chosen to be less capable than the best one available.
What the tests cannot tell you
The authors are direct about the limit. There is no ground truth for the model's real reasoning, so these are tests for specific failure modes rather than a measure of faithfulness itself. A chain can pass all four and still be unfaithful in a way nobody thought to test. They also note that every model here was RLHF-tuned, and a pretrained model might condition on its own text more strongly, since it was trained to continue text rather than to satisfy a preference model.
What we would take from the paper is the method rather than the numbers. The four interventions are cheap, need no access to weights, and give a per-task, per-model answer to whether the written reasoning is load-bearing. We would want to see them run on models trained specifically to reason at length, and on the decomposition methods the paper cites as producing more faithful chains by their own metric. If faithfulness really peaks at intermediate capability, then every new generation of larger models makes the written reasoning a less reliable window, and that is worth knowing before anyone builds oversight on top of it.
Sources
From the foundation