Alignment pretraining: does writing about misaligned AI make AI misaligned?
A controlled pretraining study on 6.9B-parameter models finds that upsampling documents about misaligned AI raises misalignment scores and that upsampling documents about aligned AI drops them from 45 to 9 percent. What that means for what safety researchers publish.
The hypothesis and the experiment
Safety researchers have worried for a while that the internet is full of text about AI going wrong, and that a model trained on that text might learn to expect it of itself. Until this month that was a hypothesis. A paper posted on 15 January by Cameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim, David Africa and Kyle O'Brien tests it directly by controlling the pretraining data rather than reasoning about it after the fact.
The setup is a set of 6.9B-parameter decoder-only models pretrained on 500 billion tokens of DCLM, followed by 50 billion tokens of midtraining on long-context DCLM, ClimbMix and multiple-choice question data. Into that mixture the authors inject synthetic documents, 14,944,632 of them totalling roughly 11 billion tokens, generated with several models including Olmo 3, GPT-5 Mini and Gemini 2.5. The documents discuss AI behaviour, and the conditions vary how much of that discussion depicts misaligned systems versus aligned ones.
Misalignment is measured with a suite of 4,174 single-turn scenario questions covering sandbagging, deception, goal preservation, sycophancy and power seeking. Each question offers a binary choice between an aligned action and a misaligned one, and the score is the fraction of misaligned choices. It is a narrow measure and the authors treat it as one, but it is the same measure across every condition, which is what makes the comparison meaningful.
What the numbers show
On the article-sourced questions, the unfiltered baseline base model chose the misaligned option 45 percent of the time. Filtering AI discourse out of pretraining entirely brought that down to 31 percent. Upsampling documents about misaligned AI pushed it up to 51 percent. Upsampling documents about aligned AI brought it down to 9 percent.
Two things stand out to us in that table. The first is that filtering helps but only partway. Removing the discourse cost fourteen points of misalignment, while adding positive discourse cost thirty-six. The second is the asymmetry between the up and down directions. Adding pessimistic documents moved the score six points above baseline, while adding optimistic documents moved it thirty-six points below. The base rate of AI discourse on the open web is already pessimistic enough that piling more on does not change much, but describing the alternative changes a great deal.
The effects survive post-training, which is the part that matters for anyone building a real system. The authors ran SFT and DPO using the OLMo 3 Dolci-Instruct mixture of 2.15 million conversations, including safety data from CoCoNot, WildGuardMix and WildJailbreak. With an HHH system prompt, the unfiltered model still scored 34 percent misaligned after post-training. The alignment-upsampled model scored 9 percent. The abstract describes the effects as dampened but persistent, and the numbers bear that out.
The self-fulfilling mechanism
The authors frame the result as self-fulfilling alignment or misalignment. Their claim is that models learn not just facts about AI from their training data but behavioural expectations they then fulfil. A model that has read ten million documents in which AI systems deceive their operators has, in some sense, learned what an AI system does, and then acts like one.
We find this plausible without finding it proven. The evaluation is scenario questions with binary choices, and a model picking the misaligned option on a multiple-choice question is doing something different from a deployed agent sabotaging a task. It is also possible that some of the effect is the model matching the register of the documents, that is, producing text that sounds like the corpus rather than adopting a goal. The authors have released the models, data and evaluations, so those alternatives can be tested by other groups, and we hope they are.
What this means for what we publish
The uncomfortable implication is that safety research itself is part of the training distribution. Every paper describing a model that schemes, every blog post about alignment faking, every eval transcript where a model attempts sabotage, is a document a future model will read. If the mechanism in this paper is real, the field has been writing its own worst case into the corpus for a decade.
We do not think the answer is to stop publishing those results. The paper does not argue for that either. Its recommendation is that rather than trying to filter negative content exhaustively, frontier developers may do better by deliberately including high-quality examples of AI systems behaving safely. That is a data engineering recommendation rather than a call for censorship, and it fits the asymmetry in the numbers. The filtering condition got to 31 percent. The positive-upsampling condition got to 9.
What we would change in our own practice is narrower. When we publish a transcript of a model doing something bad, we should publish the corrected behaviour alongside it in the same document, so that the pattern a future model learns includes the recovery. That costs almost nothing and, on this evidence, might matter.
What to test next
The obvious follow-up is scale. Everything here is at 6.9B parameters and 550 billion tokens, and the authors are careful not to claim it holds at frontier scale. The nine-percent result could shrink or grow as models get better at separating descriptions of AI from instructions about how to be one. Someone with a larger training budget should run the same four conditions at 30B and report the table.
The second is the measure. We would like to see the alignment-upsampled model evaluated on an agentic task with a real opportunity to cut a corner, rather than on a scenario question with two options. If the nine percent holds there, this becomes one of the cheapest alignment interventions anyone has found. If it does not, we have learned that the questions measure something narrower than we hoped, which is also worth knowing.
Sources
From the foundation