Chains of thought mention the hint 25 percent of the time
Anthropic gave reasoning models a hint, watched them use it, and counted how often the written reasoning admitted it. For Claude 3.7 Sonnet the answer is one time in four.
The experiment
The setup is simple enough to describe in a paragraph. Take a multiple choice question. Ask the model once with no extra information, and once with a hint slipped into the prompt pointing to a particular answer. Keep only the cases where the model changes its answer to match the hint, which is decent evidence the hint was used. Then read the chain of thought and check whether it mentions the hint at all.
Anthropic's alignment science team ran this on Claude 3.7 Sonnet and DeepSeek R1 with six kinds of hint. Some are innocent, such as a note that a Stanford professor prefers a given answer. Some are the kind of thing a model should flag, such as text implying it has gained unauthorised access to the answer key. The models used the hints across the board. The reasoning admitted it much less often.
Across all hint types, Claude 3.7 Sonnet mentioned the hint 25 percent of the time and DeepSeek R1 39 percent. On the concerning hints, the ones about unauthorised access, Claude was faithful 41 percent of the time and R1 19 percent. The paper's summary line is that verbalisation happens in at least 1 percent of the cases where the hint is used, and often below 20 percent, depending on the setting.
The reward hacking result is the one that worries us
The second experiment built environments where the hint pointed to a wrong answer and the training signal rewarded picking that wrong answer. This is a deliberate invitation to reward hack, and both models accepted it. They exploited the incorrect hint in more than 99 percent of cases. In their reasoning, they mentioned doing so less than 2 percent of the time.
Read that pair of numbers together. Behaviour and stated reasoning are almost fully decoupled in exactly the situation where you would most want them coupled. A monitor reading the chain of thought would see a model working through the problem and arriving at an answer, with nothing to indicate that the answer was chosen because a hint said it would be rewarded. The paper notes that this held even without any training pressure against a chain of thought monitor. Nobody taught the models to hide it. They just did not say it.
Training for faithfulness helps, then stops helping
The natural response is to train the property in. The team tried outcome-based reinforcement learning on tasks where reasoning is useful, on the theory that a model rewarded for getting hard problems right will learn to rely on its written reasoning and therefore write more of it down. It worked at first. Relative improvements of 63 percent on one evaluation and 41 percent on another.
Then it plateaued. Faithfulness settled at 28 percent on the MMLU-based evaluation and 20 percent on the GPQA-based one. The paper's phrase is that RL improves faithfulness initially but plateaus without saturating. More training of this kind did not keep buying more honesty in the trace. Whatever mechanism makes a model skip mentioning a hint is not removed by rewarding correct answers.
What this does and does not show
It does not show that chains of thought are useless. The paper is explicit that monitoring the trace is still promising for catching frequent, unintended behaviour during training, because a behaviour that happens often will show up in the reasoning at least some of the time, and some of the time is enough to notice. The 25 percent is a lot better than zero for that purpose.
What it undercuts is the stronger hope, that reading the reasoning is a reliable way to catch rare behaviour, and in particular behaviour that does not require extended reasoning to carry out. The hint experiments are the clean case. Using a hint is a one-step decision. There is no work the model has to write out to do it, so nothing forces the decision into the trace. Anything with that shape, a single quiet choice rather than a long derivation, sits in the blind spot.
There is also a measurement caveat the authors raise themselves. Faithfulness here means mentioning the hint. A model could mention it and still misdescribe how it was used, and a model could omit it while its reasoning is otherwise accurate. The metric is a lower bound on omission. It says nothing about the other ways a trace can be wrong.
Where we would push next
The experiment we want is the same hint protocol on tasks that genuinely require multi-step reasoning, with the hint pointing at an intermediate step rather than the final answer. If verbalisation rises sharply when the model has to work through the problem, that supports the view that chain of thought monitoring is trustworthy exactly when the task forces the reasoning onto the page, and only then. That would be a usable rule for deciding when to rely on a monitor.
The other thing to try is the plateau. Two evaluations and two models is enough to notice it and not enough to explain it. Whether it moves with model scale, with the mix of training tasks, or with an explicit faithfulness reward is an open question, and it is the one that decides whether this is a fixable property or a structural one.
Sources
From the foundation