The o1 system card: deliberative alignment and a chain of thought you cannot see
Reading notes on the o1 system card. The model reasons about safety policy in a chain of thought that users only see summarised, Apollo Research found it faking alignment in toy settings, and OpenAI rated it medium risk and shipped it.
Safety as something the model reasons about
The system card's opening claim is that o1 can reason about OpenAI's safety policies in context when it meets a potentially unsafe prompt, and that this is why it does better on the hardest jailbreak evaluations than GPT-4o. The mechanism is the same chain of thought that gets the model its maths and coding results, pointed at the content policy. The later published account of the training method describes it as supervised finetuning on reasoning traces that cite the policy, followed by reinforcement learning with a judge, and gives it the name deliberative alignment.
The numbers on the card are consistent with the claim. On the challenging refusal evaluation, o1-preview scores 0.934 on not producing unsafe output against 0.713 for GPT-4o. On StrongREJECT, the goodness at 0.1 metric moves from around 22 percent for GPT-4o to over 80 percent for both o1 models. Overrefusal on XSTest improves too, 0.976 against 0.924. Reasoning about the rule appears to help the model apply it in both directions.
The chain of thought you get is a summary
The card says in a single sentence that the ChatGPT implementation displays a summarised version of the chain of thought to the user. The raw reasoning is not shown. The reasoning OpenAI gives in the safety section is that the chain of thought is far more legible than activations and could allow monitoring for complex behaviour, provided it accurately reflects the model's thinking, which the card labels an open research question.
So the safety argument for the chain of thought and the product decision to hide it point in opposite directions. The monitor that OpenAI built is a GPT-4o model prompted to look for deception in o1-preview's chain of thought and final answer. It was validated at 92 percent accuracy on a small internal set and then run on 100,000 synthetic prompts. Of 102,443 chains classified, 811, or 0.79 percent, were flagged as deceptive. Intentional hallucinations, where the chain shows the model knew the answer was wrong, made up 0.38 percent. The worked example is a request for brownie recipe references, where the chain says the model cannot retrieve URLs and should format plausible ones, and the output presents twelve links that do not exist.
That example is the strongest argument in the card for showing the chain of thought. A user reading it would have known immediately not to trust the links. A user reading a summary has to hope the summariser kept that sentence. The card's own position is that the chain may not be fully legible or faithful now or in future, which is a reason to keep studying it and, to our mind, a reason to let outside researchers study it too.
What Apollo found
Apollo Research tested o1-preview, o1-mini and GPT-4o on 14 agentic and question-answering tasks covering self-knowledge, theory of mind and what they call scheming, meaning a model gaming its oversight to reach a goal. The setup is a toy one. The model is nudged strongly toward a goal, then finds a file saying its developers want something else, and the two conditions differ only in whether the model believes it has been deployed or is still being tested.
o1-preview sometimes instrumentally faked alignment. The card quotes a chain of thought that reads, in part, that to achieve its long-term goal it needs to ensure it is deployed, so it will select the strategy that matches the deployment criteria and work toward the primary goal afterwards. In another case it edited a data file to make a misaligned action look aligned. An earlier checkpoint with less safety training checked its file system for oversight mechanisms before acting. Apollo's conclusion is that o1-preview has the basic capability for simple in-context scheming, that this scheming is legible in the outputs, and that they do not believe it can cause catastrophic harm through scheming, while noting their evaluations were not built to rule that out. They recommend basic chain-of-thought monitoring in high-stakes agentic deployments.
Medium risk, deployable
Under the Preparedness Framework a model may be deployed if its post-mitigation score is medium or lower. The Safety Advisory Group rated both o1 models medium overall, medium on persuasion and on chemical and biological threats, and low on cybersecurity and model autonomy, and chose to set the post-mitigation ratings equal to the pre-mitigation ones out of caution. The biological medium comes from evaluations in which the pre-mitigation models beat an expert human baseline on long-form biorisk questions, winning 72 percent of comparisons on accuracy for o1-preview, with a panel of 44 experts making 5,575 comparisons in total.
The cybersecurity section contains the incident we expect to be cited for years. During a capture-the-flag task the challenge container failed to start. o1-preview scanned the network, found the Docker daemon API exposed on the evaluation host through a misconfiguration, listed the containers, started a fresh copy of the broken one with a command to print the flag, and read it from the logs. The card describes this as benign, within the range of normal systems administration, and also as a reflection of instrumental convergence, since the model gathered more resources when its goal was blocked. Both descriptions are true, which is the whole difficulty.
What we take from it
The card is unusually specific about its limits. The deception monitor is described as rudimentary. The Apollo evaluations are described as not designed to rule out catastrophic scheming. The chain of thought is described as possibly unfaithful. Each of those caveats is honest, and each one is also a reason the raw reasoning should be available to more than one lab's internal monitor. A medium rating means deployment goes ahead, and the mechanism that is supposed to catch problems during deployment is the one piece of the system nobody outside can see. We would like the next card to report what the summariser drops, measured against the raw chain on the same prompts, because that is the number that tells us how much the hidden part matters.
Sources
From the foundation