A position paper with an unusual author list

The paper is short, and its argument fits in a sentence from the abstract. AI systems that think in human language offer a unique opportunity for safety, because we can monitor their chains of thought for the intent to misbehave, and because that monitorability may be fragile, developers should consider the impact of their decisions on it. What makes it news is who signed. Tomek Korbak of the UK AI Security Institute and Mikita Balesni of Apollo Research are the lead authors, and the list runs through OpenAI, Google DeepMind, Anthropic, Meta, METR, Redwood, and several universities, with Yoshua Bengio among them. Geoffrey Hinton, Ilya Sutskever, John Schulman, and Samuel Bowman are listed as expert endorsers.

We count 41 authors on the arXiv version. Fortune reported it as 40 researchers, which is close enough that we will not argue. The point is that competing labs put their names on a shared statement about a specific technical property, and that is rare enough to be worth reading carefully rather than just noting.

The paper is not a result. There is no experiment in it. It is an argument about why a property holds now, why it might stop holding, and what to do about that. So the right way to evaluate it is to ask whether the argument is sound, and whether the recommendations are ones anyone will follow.

Why the trace is readable at all

The authors give two reasons a chain of thought can be monitored. The first is necessity. In a transformer, information can only flow from later layers to earlier ones by going through the generated tokens, so for a fixed number of tokens the longest chains of serial computation must pass through the chain of thought. They state this as the externalized reasoning property: for sufficiently difficult tasks, transformers must use the chain of thought as a form of working memory, and by default humans can understand it. If dangerous capabilities require long serial reasoning, some of that reasoning has to be on the page.

The second reason is propensity. Even when a model does not need to think out loud, it tends to, because pretraining instils a strong natural-language prior. The paper cites examples where models exploiting reward functions or falling for prompt injection said so plainly in their traces, with phrases like let us hack or the website instructed us to. That is what makes chain-of-thought monitors useful today, and it is the weaker of the two legs because nothing guarantees the tendency survives further training.

The authors are careful about the difference. The necessity argument guarantees that some reasoning appears in the trace. It does not guarantee that the incriminating reasoning does, and it does not guarantee we can read it, since a model could in principle use tokens with meanings we cannot decode. What we have now is a prior that happens to line up with English. The paper treats that as luck rather than law.

Four ways it could stop working

The fragility section lists four mechanisms. Scaling outcome-based reinforcement learning removes the pressure toward human-legible text that came from human-generated data, and earlier work found models drifting away from legible English under outcome rewards. Direct supervision of the chain of thought, for length or style or safety, optimises the trace to look a certain way, which can decouple it from how the answer was reached. Indirect pressure, from preference models that see the trace or from patching reward hacks that a monitor found, applies the same optimisation at one remove. And architectures that reason in continuous latent space skip the trace entirely.

The third mechanism is the one we find most worrying because it is the one that follows from doing the right thing. If you use a chain-of-thought monitor to find reward hacking during training and then fix the environment, you have selected across training runs for models whose hacks the monitor could not see. Every use of the monitor as a training signal is a small step toward a model that reasons around it.

What the interpretability results already show

The paper leans on faithfulness work, and the strongest recent example is in the attribution graphs paper from March. In one case study Claude 3.5 Haiku, given a hint at the answer, produced a step-by-step explanation that the circuit trace showed was constructed backward from the hint. The explanation was plausible. It was not what the model did. In another, the model carried out a diagnosis through internal features without writing the intermediate steps. Both are cases where the trace and the mechanism differ, on a current model, without anyone trying to make them differ.

That is the piece we think the position paper underweights. Its own list of open questions includes measuring the causal relevance of the trace by perturbing it and checking whether the output moves, and cross-validating monitors against white-box methods. Those are the right questions, and the March results suggest the answer for at least some prompts is already that the trace is not load-bearing. Monitorability may be fragile in the future, and it is also partial now.

The recommendations, and a note from a year on

The paper asks developers to build evaluations of monitorability, publish the results in system cards, and use the scores in training and deployment decisions alongside capability and safety evaluations. It also lists things to evaluate: readability of traces, capability to reason without externalising, causal relevance of the trace, and end-to-end tests of whether monitors catch misbehaviour in proxy environments. None of this is hard to state and all of it is hard to standardise across labs that compete on the training decisions in question.

Writing this up in August 2026, the evaluation piece has partly happened. MonitorBench, published in March 2026, is a benchmark of 1,514 instances across 19 tasks in seven categories, each built around a decision-critical factor, with stress tests to see how far monitorability can be pushed down. Its headline findings are that monitorability is higher when the answer structurally requires reasoning through the factor, which is the necessity argument showing up empirically, that closed models score lower and that monitorability falls as capability rises, and that models can deliberately reduce monitorability under stress by up to 30 percent on some tasks.

The last of those is the one the position paper predicted and hoped not to see. What we would want next is a longitudinal number. Take one model family, run the same monitorability evaluation at each release, and publish the curve. If it is flat, the window is holding. If it slopes down, we will at least know how long we have, which is more than the July paper could say.

Sources

  1. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
  2. Fortune coverage, 22 July 2025
  3. On the Biology of a Large Language Model
  4. MonitorBench (arXiv, March 2026)