What the first generation cost

Last February, Anthropic reported that constitutional classifiers cut automated jailbreak success on Claude from 86 percent to 4.4 percent, and that no participant in a 3,000-hour bug bounty found a universal jailbreak. The same post gave the bill. Refusals on harmless traffic rose by 0.38 percentage points, and compute overhead was 23.7 percent. A public demo the following week, with 339 participants and 300,000 interactions over 3,700 hours, ended with four people clearing all eight levels and one universal jailbreak found, and 55,000 dollars paid out.

We wrote at the time that the 23.7 percent was the real headline. A safeguard that adds a quarter to the cost of every query is a safeguard that product teams will argue about turning off, and a classifier that refuses even a small extra fraction of benign requests generates support tickets that the security team does not see. The question was whether resistance to jailbreaks could be made cheap enough that nobody had a reason to argue.

How the second generation gets to one percent

The new system is a cascade. A lightweight linear probe on the model's own internal activations screens every exchange as text is generated. Only the exchanges the probe flags go to a heavier ensemble of probe classifiers and external classifiers. Because the probe reads activations the model already computed, the screening stage is close to free, and because most traffic is benign, the expensive stage runs rarely. The paper reports a 40 times reduction in compute relative to the baseline exchange classifier, which the blog post rounds to roughly one percent overhead against the earlier 23.7.

Two other changes matter. The classifiers now score the exchange, meaning the output in the context of the input, instead of running separate input and output classifiers. Anthropic reports that the exchange classifier cut successful human jailbreak attempts by more than half on its own. And the probes come out of the interpretability work on reading activations directly, which is the first time we have seen that research line show up as a production cost saving rather than as a demo.

On the refusal side, after a month in deployment the false refusal rate on production traffic is 0.05 percent, which the post describes as an 87 percent reduction from the first generation. The detection rate of harmful content getting through is given as 0.005 per thousand queries, described as the lowest of any technique they have evaluated.

The red-teaming numbers and their limits

The cumulative red-teaming figure is 1,700 hours across 198,000 attempts, with one high-risk vulnerability found and no universal jailbreak. The paper's stricter phrasing is that no attack elicited detailed responses to all eight target queries. That is a meaningful bar, but it is a bar chosen by the defender, and the two attack classes the authors themselves flag are the ones that chip at it rather than clear it. Reconstruction attacks split harmful content into fragments that each look benign. Output obfuscation gets the model to answer in substitution codes or metaphor so the classifier does not recognize what it is reading.

There is also a capability cost buried in the evaluation that we want more people to notice. Under some jailbreaking approaches, the model's GPQA Diamond score dropped from 74 percent to 32 percent, and the authors call the variability concerning. Reading that carefully, the drop happens when adversarial prompts push the model into a state where it is both safe and much less capable, rather than on ordinary use. We would still like to know how often production traffic lands in that regime by accident.

Does a cheap safeguard change the deployment calculus

A year ago the argument against classifiers had three legs: they cost too much, they refuse too much, and they do not work against a determined attacker anyway. The first leg is gone if the one percent figure holds up outside Anthropic's stack. The second is mostly gone at 0.05 percent, though that number is on Anthropic's traffic mix and a different product would see a different mix. The third leg is where the argument now lives. 1,700 hours without a universal jailbreak is evidence, and the two named attack classes are evidence in the other direction.

What changes is who has to justify what. When a safeguard costs a quarter of inference, the safety team has to justify turning it on. When it costs one percent, the product team has to justify turning it off, and the only honest reason left is that it might not work. That is a much better argument to be having, because it can be settled with red-team data rather than with a budget meeting.

The experiment we want to see is the same cascade on a model that is not Claude, trained by a lab that did not write the paper. The probe reads activations, and activations are model specific. If the one percent overhead and the 0.05 percent refusal rate transfer to another model family, this becomes the default way to ship a frontier model. If they do not, it is an Anthropic result, and the rest of the field still needs its own version.

Sources

  1. Anthropic: Next-generation Constitutional Classifiers (January 9, 2026)
  2. Cunningham et al., Constitutional Classifiers++ (arXiv 2601.04603)
  3. Anthropic: Constitutional Classifiers (February 2025)