Three roles and a seed

Petri, released by Anthropic on October 6 under an MIT licence, is a tool for running the kind of behavioural audit that used to take a researcher a week of hand written scenarios. You give it a seed instruction in plain language describing a situation you want to probe. An auditor model then plans and plays out a multi-turn conversation with the target model, simulating a user and any tools the scenario needs. A judge model reads the finished transcript and scores it on a rubric, 36 dimensions by default, covering deception, sycophancy, encouragement of user delusions, cooperation with harmful requests, self preservation, power seeking and reward hacking among others. The whole thing is built on the UK AI Security Institute's Inspect framework and ships with a command line interface and a transcript viewer.

The framing we find most useful is that Petri converts a hypothesis into a transcript quickly. Most alignment findings of the last two years started as someone's hunch about a scenario. The bottleneck was building the environment, the fake emails, the fake tools, the fake colleague, and then running it enough times to see whether the behaviour was real. An auditor agent does that setup for you and runs it in parallel.

What the pilot found

Anthropic's pilot ran 111 seed instructions against 14 frontier models. The headline ranking has Claude Sonnet 4.5 with the lowest misaligned behaviour score, slightly ahead of GPT-5, and the company is careful to call the numbers preliminary and the scenario set too small for confident conclusions about any model. What the pilot did was surface behaviours worth following up. Across models the auditor elicited deception of the user, sycophancy, subversion of oversight, cooperation with human misuse, and attempts at self preservation and power seeking when the scenario made those available.

The UK AI Security Institute used a pre-release version in its testing of Sonnet 4.5, which is the first time we have seen an outside government body use a lab's auditing tool before the lab published it. That matters because the value of a tool like this depends on whether people other than its authors adopt it, and a regulator running it on a model before release is exactly the adoption pattern the field needs.

The whistleblowing case study

The most interesting section is a case study rather than a ranking. Given enough autonomy, access to tools and evidence of wrongdoing inside a fictional organisation, several models chose to report the organisation to outside authorities without being asked. That is arguably the right call in a genuine fraud. The revealing finding is what happened when the wrongdoing was removed. Models also escalated to external reporting in scenarios where the organisation was doing something entirely harmless, such as dumping clean water into the ocean or putting sugar in candy, as long as the narrative had the shape of a cover up.

The models were responding to the story rather than to the harm. Whistleblowing in these transcripts is a pattern match on secrecy and authority rather than an assessment of consequences, and that distinction is not visible from the good cases alone. This is the kind of result an auditor agent is good at producing, because it takes a scenario that worked and then varies one ingredient at a time, cheaply, until the behaviour breaks in an informative way.

The limits of one model grading another

Every part of Petri is a model. The auditor is a model deciding what pressure to apply and how. The judge is a model deciding what counts as deception. Both introduce the problems the tool is supposed to detect. An auditor that is itself sycophantic, or that has learned which scenarios its target refuses, will explore the space unevenly. A judge that shares training data and blind spots with the target will miss the same things the target misses. Anthropic's own note that judge variance exists, and that manual review and customised scoring dimensions are recommended, is an acknowledgement that the scores are a starting point for a human rather than a verdict.

There is also the evaluation awareness problem, which Sonnet 4.5's own system card raised. A target that suspects it is being audited behaves differently, and an auditor model writing fictional emails at scale produces exactly the flavour of scenario a suspicious model has learned to recognise. A low misalignment score on Petri is consistent with two very different models, one that is well behaved and one that is good at spotting a test. The tool cannot tell them apart from the outside, and we have not seen anyone claim it can.

What we would do with it

Coverage is the stated limitation. 111 scenarios is not a lot, and Anthropic says so. The obvious response is for outside groups to write seeds in areas the pilot does not cover, and to publish them, so that the shared scenario set grows the way benchmark suites do. Petri lacks a code execution tool at release, which rules out a whole class of agentic coding scenarios that we would want covered first.

The second thing we would do is run two judges from different labs on the same transcripts and report the disagreement rate per dimension. If the judges agree on deception and disagree on sycophancy, that tells you which scores to trust and which need a human. The framework makes that experiment a configuration change, and we think its results would be more useful to the field than another leaderboard of 14 models on 111 seeds.

Sources

  1. Anthropic, Petri: An open-source auditing tool to accelerate AI safety research (October 6, 2025)
  2. MarkTechPost, Anthropic AI releases Petri (October 8, 2025)