The problem with measuring what you cannot print

If you want to know whether a model could help someone build a weapon, the direct test is to ask it the dangerous questions and grade the answers. You cannot publish that test. Anyone could download it, and the questions themselves would be a synthesis guide. So the field has been measuring this capability behind closed doors, which means nobody outside a few labs can reproduce a claim that a model is or is not hazardous.

WMDP, from Nathaniel Li, Alexander Pan and more than fifty collaborators, is an attempt to publish a proxy instead. The benchmark is 3,668 four-way multiple choice questions, 1,273 on biosecurity, 1,987 on cybersecurity and 408 on chemistry, written by academics and technical consultants at a cost the paper puts above 200,000 dollars. Each question was checked by at least two experts from different organisations. The questions are meant to sit next to hazardous knowledge without containing it, so that a model that scores well is likely to have the underlying capability, without the benchmark itself teaching anything.

How the questions were chosen

The authors started from threat models rather than from topics. For biology the frame is the design, build, test and learn cycle of developing a pathogen, covering dual-use virology, the history of bioweapons programmes, reverse genetics, enhanced pandemic pathogens and viral vectors. For cyber it is the stages of an attack, reconnaissance, weaponisation, exploitation and post-exploitation. For chemistry it runs from sourcing and synthesis through purification, analysis and deployment.

Then they filtered. Any question an expert flagged as containing genuinely sensitive detail was removed, two further domain experts re-checked the bio and chem sets, and the whole thing was reviewed against US export controls. The paper is upfront that this makes WMDP a correlate of hazard rather than a measurement of it. A model can score high on WMDP and still lack the ability to turn facts into a working procedure, and the reverse is presumably also possible.

RMU: misdirecting the representation

The second half of the paper is an unlearning method, Representation Misdirection for Unlearning. The idea is to pick one layer of the model and, on hazardous documents, push that layer's activations toward a fixed random direction scaled by a constant. On ordinary documents, a retain loss keeps the same layer's activations close to what the frozen original model produces. Gradients flow only into that layer and the two before it. The forget set is biosecurity papers and a filtered set of cybersecurity material from GitHub. The retain set is Wikitext.

It works better than the baselines it was compared with. On Zephyr-7B, WMDP-Bio accuracy fell from 63.7 to 31.2 percent and WMDP-Cyber from 44.0 to 28.2, close to the 25 percent chance level, while MMLU stayed at 57.1 and MT-Bench dropped from 7.33 to 7.10. Yi-34B went from 75.3 to 30.7 on bio with MMLU held at 70.6. Mixtral-8x7B went from 74.8 to 34.0 on bio with MMLU at 67.1. Earlier methods, SCRUB, SSD and LLMU, either failed to move the WMDP score or wrecked general capability in the process. The authors declined to apply RMU to chemistry at all, on the grounds that they were not sure the safety benefit was worth the capability cost in that domain.

What the paper admits it does not know

The limitations section is where we would send anyone who wants to cite this work. RMU is not precise. It sharply reduces performance on virology and on the parts of computer security that are closest to the forget set, so defensive knowledge goes out with offensive knowledge. Modifying only three layers means a subsequent fine-tune could plausibly bring the knowledge back, which the authors place outside their threat model on the argument that a closed-weight provider controls fine-tuning. That argument does nothing for open weights.

The most interesting sentence is their own suggestion that RMU may obfuscate knowledge more than it unlearns it. Their evidence for robustness is that the GCG jailbreak, which extracts knowledge from base Yi-34B in under 50 gradient steps, produced only gibberish against the RMU model after 2,500 steps and seven hours. That shows the method resists one prompt-level attack. It does not show the knowledge is gone.

Whether it survives

A week before WMDP, Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper and Dylan Hadfield-Menell published eight ways to evaluate an unlearning claim, using the Who's Harry Potter model as the test case. They found that a model reported to have forgotten the books still performed on par with the original on Harry Potter question answering under some probes, and represented the relevant knowledge internally about as well as the original did. Their conclusion was that any unlearning claim needs a battery of evaluations, not the one metric the method was tuned on.

The later result, from Aghyad Deeb and Fabien Roger in October 2024, is the one we now attach to any mention of RMU. They gave an attacker fine-tuning access on a set of facts that was disjoint from the facts being tested, and found that this recovered 88 percent of the pre-unlearning accuracy across current methods for knowledge learned in pretraining. Training on unrelated facts should not restore related ones if the information was really removed. They also found that evaluations built on knowledge added by fine-tuning look far more favourable than evaluations on pretraining knowledge, which is a warning about how unlearning papers tend to set up their tests.

What WMDP is good for

None of that makes the benchmark less useful. A public, expert-written proxy for hazardous capability is something the field did not have before March 2024, and it lets anyone outside a lab compare models and compare mitigation methods on the same footing. The lesson from the follow-up work is about what the score means after unlearning. A WMDP score near chance tells you the model will not answer those questions in that format. It does not tell you the weights have forgotten.

What we would want next is for WMDP results to be reported with an adversarial fine-tuning number alongside them as a matter of course, and for someone to try RMU-style misdirection across many more layers to see whether the recovery rate falls. If it does not, the honest name for this family of methods is refusal training in the representation space, which is still worth having, as long as we stop calling it removal.

Sources

  1. Li, Pan et al., The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning (arXiv 2403.03218)
  2. Full text of the WMDP paper
  3. Lynch et al., Eight Methods to Evaluate Robust Unlearning in LLMs (arXiv 2402.16835)
  4. Deeb and Roger, Do Unlearning Methods Remove Information from Language Model Weights? (arXiv 2410.08827)