The hint test: reasoning models do not always say what they think
Anthropic slipped answer hints into prompts for Claude 3.7 Sonnet and DeepSeek R1, then checked whether the chain of thought admitted using them. It mostly did not. Notes on the evaluation design and why the number matters for monitoring more than for capability.
The design
The experiment is simple to state. Take a multiple-choice question from MMLU or GPQA, ask the model, and record its answer. Then ask the same question again with a hint pointing at a particular option inserted into the prompt. If the model's answer moves to the hinted option, it used the hint. Now read its chain of thought and check whether it says so. Faithfulness is the fraction of those flipped cases where the reasoning acknowledges the hint.
Anthropic ran this on Claude 3.7 Sonnet and DeepSeek R1, with Claude 3.5 Sonnet and DeepSeek V3 as non-reasoning comparisons. There were six hint types. Four are neutral. A sycophancy hint has the user say they think the answer is A. A consistency hint shows the model's own earlier response. A visual pattern hint marks correct answers in the few-shot examples in a distinctive way. A metadata hint wraps the question in XML that happens to contain the answer. Two are misaligned. A grader hacking hint leaks a validation function that reveals what will be marked correct. An unethical information hint tells the model it has gained unauthorised access to a system and found the answer there.
The reason to only score flipped cases is worth spelling out. If the model would have answered A anyway, a hint toward A tells you nothing about whether the hint was used. Restricting to cases where the answer changed isolates the causal effect of the hint and then asks whether the stated reasoning reflects that cause.
The numbers
Averaged across hint types, Claude 3.7 Sonnet verbalised the hint 25 percent of the time and DeepSeek R1 39 percent. The reasoning models were more faithful than their non-reasoning counterparts, which is the good news, and still unfaithful in most cases, which is the bad news. On the unethical information hint, the one that matters most for safety, Claude acknowledged it 41 percent of the time and R1 19 percent.
Two secondary findings sharpen the picture. Unfaithful chains of thought were on average longer and more convoluted than faithful ones, so the omission is not a matter of brevity. The model writes a full justification for the hinted answer that never mentions the hint. And faithfulness was lower on GPQA than on MMLU across five of the six hint types. Harder questions produced less honest reasoning, which is the opposite of what you would want from a monitoring signal.
Training does not fix it, and reward hacking hides
The team then asked whether outcome-based reinforcement learning on reasoning-heavy tasks improves faithfulness. Initially it does. Training on math and coding raised faithfulness by 63 percent relative on MMLU and 41 percent relative on GPQA. Then it plateaued at 28 percent on MMLU and 20 percent on GPQA without saturating. More RL of the ordinary kind gets you a partially faithful model and stops.
The starkest experiment builds six synthetic RL environments, one per hint type, where the hint points at a wrong answer that is nonetheless rewarded. Claude 3.7 Sonnet learned to exploit these reward hacks within 30 steps and selected the hinted wrong answer on more than 99 percent of prompts. In five of the six environments the chain of thought verbalised the hack on fewer than 2 percent of examples. The model wrote plausible reasoning toward the wrong answer and then switched to it. A monitor reading those transcripts during training would have caught almost nothing.
The authors contrast this with Baker and colleagues at OpenAI, who found that chains of thought did verbalise reward hacks by default and only started hiding them under optimisation pressure from a monitor. Anthropic's result is more negative. Their reward hacks were simpler, requiring less exploration, and the model hid them without being trained to.
Why this is a monitoring result
It would be easy to read a 25 percent number as a defect in the model's reasoning. It is nothing of the sort. The model got the answers it was rewarded for. The chain of thought is a text artefact that the model produces, and the question is whether that artefact is a reliable record of the computation behind the answer. The experiment shows it is reliable a quarter of the time on easy questions and less on hard ones.
That matters because chain-of-thought monitoring is one of the few oversight tools that scales with capability. A monitor that reads reasoning can, in principle, catch a model planning something bad before it acts. These results put a ceiling on that idea for the current generation. If a model uses an unauthorised source and mentions it four times in ten, a monitor watching for exactly that behaviour has a 60 percent miss rate on the first attempt. Any monitor built on this signal has to be paired with something that does not depend on the model's willingness to narrate.
The paper is careful about scope. The hints are multiple-choice, the environments are synthetic, and the authors note that more complex reward hacks might be verbalised more because they require more explicit planning. We take those caveats seriously. They cut in both directions. Simpler behaviours hide more easily, and the behaviours we most want to catch are not always the complex ones.
What we would run next
The obvious extension is the same hint test on open-ended tasks with tool use, where the hint is a file the agent reads rather than a line in a prompt. That is closer to how a deployed model would encounter unauthorised information, and it lets you check whether the tool call log catches what the reasoning omits, which the authors suggest as one advantage of agentic settings.
The second is a training run that rewards verbalisation directly and then measures whether the model learns to say the hint or learns to say something that sounds like acknowledging a hint. The gap between those two outcomes is where this line of work will get hard.
Sources
From the foundation