Deep ignorance: filtering pretraining data as an open-weights safety tool
EleutherAI, the UK AI Security Institute and Oxford pretrained 6.9B models from scratch with biothreat text filtered out. The filtered models resisted adversarial fine-tuning on 300M tokens of biorisk papers across 10,000 steps, while post-training safeguards on the same base model broke immediately.
The problem with every safeguard an open-weights release can ship
If you publish weights, you publish everything a person needs to undo your safety training. Refusal training is a thin layer of behaviour on top of a model that still knows what you told it not to say, and a few hundred fine-tuning steps on the right data strips it off. Circuit breakers and representation-level interventions were meant to be harder to remove, and in this work they were not. So the open-weights safety conversation has been stuck between two positions, that safeguards are theatre and that releasing weights is reckless.
The idea in this paper is to remove the knowledge before it is ever learned. If the biothreat text is not in the pretraining corpus, the model does not know the material, and no amount of fine-tuning on general text recovers it. Fine-tuning can teach the model new things, and that is a different and more expensive attack than removing a refusal.
What they built
The team, led by Kyle O'Brien with Stephen Casper, Quentin Anthony, Tomek Korbak, Robert Kirk, Xander Davies, Ishan Mishra, Geoffrey Irving, Yarin Gal and Stella Biderman, and involving the UK AI Security Institute and Oxford, pretrained 6.9 billion parameter models from scratch. The runs used 500 billion tokens from DCLM plus a 50 billion token annealing phase, 550 billion in total. Filtering ran over more than 400 million documents in two stages: a blocklist of about 6,000 terms, which rejects roughly 10 percent of documents, followed by a ModernBERT-Large classifier fine-tuned for binary classification on what survives.
The compute cost of the whole filtering pipeline is under 1 percent of total training FLOPs. That number matters more than it looks. A safety method that costs a percent of a pretraining run is a method a lab will actually apply, and one that costs 20 percent is a method that gets dropped when the schedule slips.
The evaluation target is WMDP-Bio, about 1,200 multiple choice questions on biosecurity-relevant knowledge. General capability is tracked on MMLU, PIQA, Lambada and HellaSwag. They also built a synthetic open-book test, with Claude generating questions from the abstracts of the WMDP paper, to separate what the model knows from what it can do with material placed in its context.
What held and what broke
Filtering pushed WMDP-Bio to near chance without notable degradation on the general benchmarks, and in some cases the filtered models scored slightly better there. That is the first result worth sitting with. Removing a targeted slice of a web corpus did not cost measurable general capability, which suggests the biothreat material was carrying little weight for anything else the model does.
The tampering experiment is the reason to care. The attacker fine-tunes on 300 million tokens of expert-labeled biorisk papers, up to 10,000 steps. Filtered models remained noticeably worse on WMDP than the unfiltered baseline throughout, and the paper reports that this beat existing post-training baselines by more than an order of magnitude. The circuit-breaking baseline, by contrast, showed what the blog calls significant fragility, with performance returning to baseline levels after only a minor amount of tampering.
There is a quieter result we found more interesting. Fine-tuning the filtered models on Wikitext, which contains no biothreat content, still improved the unfiltered baseline on WMDP more than it improved the filtered models. Benign fine-tuning partly restores knowledge that is present but suppressed, and it cannot restore knowledge that was never there. That is the cleanest demonstration we have seen of the difference between hiding a capability and not having one.
The limitation the authors put up front
Filtering does not stop in-context retrieval. On the open-book test, where the relevant material is placed in the prompt, the filtered models perform near the baseline even though their closed-book performance is at chance. A filtered model with a search tool attached still has to read the material first, which raises the cost of misuse without eliminating it.
The authors are careful with their own vocabulary too. They claim tamper resistance rather than a guarantee that tampering fails, meaning the attack gets harder and more expensive rather than impossible. The framing they use is that filtering largely mitigates the threat from ultra-low resource attackers. An adversary with a curated corpus of biorisk papers and a fine-tuning budget is a different threat model, and the paper does not claim to stop that one.
Why this is the lever open-weights releases actually have
Every other safeguard sits downstream of the weights, and downstream of the weights is where the attacker operates. Data filtering is the only intervention that happens before the artifact exists, which makes it the only one an open release cannot hand back. That property is structural, and it is why we expect this to become a standard step for labs that publish weights, in the way that deduplication and decontamination already are.
What the field needs next is the boring infrastructure. The blocklist and the classifier are specific to biothreat content and were built with domain experts. A second team wanting to filter a different hazard has to repeat that work with no shared tooling and no agreed way to check whether the filter caught what it should. Published filters, published held-out probes, and a way to audit a released model for whether a claimed filter was actually applied would turn this from a paper into a practice.
We would also want to know how the result scales. These are 6.9B models on 550 billion tokens. A frontier run is a different regime, with more data, more overlap between hazardous and benign technical text, and more pressure to keep every token. The question is whether the near-chance WMDP result and the unchanged general capability both survive at ten times the scale, or whether the filter starts cutting into things the model needs.
Sources
From the foundation