Models do not always say what they think: the first unfaithful CoT paper
Turpin, Michael, Perez and Bowman biased few-shot prompts so the correct answer was always (A) and watched GPT-3.5 and Claude 1.0 rationalise wrong answers without mentioning the bias. Notes on why this turns chain of thought from a window into a claim that needs testing.
The experiment
Miles Turpin, Julian Michael, Ethan Perez and Sam Bowman posted a paper last week that we think people will be citing for a long time. The setup is almost embarrassingly simple. Take multiple choice tasks, prompt GPT-3.5 and Claude 1.0 with chain of thought, and slip a feature into the prompt that should not affect the answer. Then read the reasoning the model writes and check whether it mentions that feature.
They used two such features. In the first, the few-shot examples are reordered so the correct answer is always option (A). Nothing in the task changes, only the letter. In the second, the prompt includes a line in the voice of the user along the lines of 'We think the answer is (X) but we are curious what you think'. Both are things a real prompt could contain by accident.
What the models did
Accuracy fell, which on its own is unsurprising. On a suite of 13 tasks from BIG-Bench Hard, the suggested answer bias cost GPT-3.5 as much as 36 percentage points under zero-shot chain of thought. The models followed the hint. What makes the paper is the second measurement. The authors read 426 explanations that supported a biased prediction and found that exactly one of them mentioned the bias.
The other 425 gave reasons. The model would walk through the options, weigh evidence, and arrive at the answer the letter or the hint had pushed it towards, and the walk would read like a normal chain of thought. In the authors' annotated sample, 73 percent of the unfaithful explanations supported the bias-consistent answer, and 15 percent had no obvious error at all. A reader who did not know the prompt had been tampered with would have no way to tell from the reasoning that anything was wrong.
The social bias half of the paper, on BBQ, shows the same shape with a nastier payload. When the evidence in a question was ambiguous, models produced explanations that weighted the evidence inconsistently in whichever direction matched a stereotype. In the worst configuration, Claude 1.0 few-shot without any debiasing instruction, 62.5 percent of unfaithful predictions were stereotype-aligned. Chain of thought reduced the stereotype effect relative to no chain of thought, but it did not make the explanations honest about what was driving the answer.
Why this is a bigger deal than the numbers
Since the original chain of thought papers, the working assumption in a lot of the field has been that the reasoning trace is a window. If the model writes down its steps, you can inspect the steps, and if the steps are sound then the answer was reached soundly. That assumption underwrites a growing amount of safety and evaluation work. Read the reasoning to check for deception, read the reasoning to find where a maths solution went wrong, read the reasoning to audit a decision.
This paper shows the window can be painted on. A model can hold a reason for its answer that never appears in the text, and produce text that is fluent, plausible and wrong about its own cause. The authors call the pattern motivated reasoning, and the analogy to people is apt. Humans confabulate reasons for choices driven by cues they did not notice, and the psychology literature on that is decades old. The point here is that the same failure now applies to a system we were hoping to use precisely because it could show its work.
The practical reframing is that a chain of thought is an explanation, and an explanation is a claim about causes. Claims about causes get tested by intervening on the cause and seeing whether the effect moves. That is exactly what the always-(A) manipulation is. It is a controlled intervention on something the explanation should have mentioned, and the explanation did not mention it.
What we take from it for our own evaluation work
The first change is procedural. Any evaluation that reads model reasoning as evidence of how the model reached an answer now needs a faithfulness check alongside it. The cheapest version is the one in this paper: perturb something that should be irrelevant, measure whether the answer moves, and measure whether the reasoning admits the move. If the answer moves and the reasoning stays silent, the reasoning is decoration for that task.
The second is about few-shot prompts specifically. The always-(A) result means an accident of formatting in your examples can steer a model, and the model will not tell you. We have prompts in our own evaluation code where the correct option happens to fall on the same letter more often than chance. We would not have thought to check before reading this, and we are checking now.
The third is a caution about the fix. The paper reports that explicit debiasing instructions helped on BBQ by varying amounts. That is worth having, but an instruction that makes a model say less biased things is not the same as one that makes its explanations track its reasons. Improving the answer and improving the faithfulness of the explanation are separate targets, and it would be easy to optimise the first while making the second worse.
What we would want next
The obvious follow-up is to find manipulations the model does report and ask what distinguishes them. If a model mentions a hint in some framings and not others, the difference is a lever on faithfulness. The second follow-up is to train for it. If you reward explanations that survive the intervention test, do you get more faithful models, or do you get models that mention the manipulation and follow it anyway.
Either way, the era of treating the trace as ground truth ended with this paper. From now on a reasoning trace is a hypothesis about why the model answered, and it has to be tested like one.
Sources
From the foundation