What was run

On August 27 OpenAI and Anthropic each published findings from running their own alignment evaluations on the other lab's models. The work was done in June and early July. Anthropic evaluated GPT-4o, GPT-4.1, o3 and o4-mini, and compared them against Claude Opus 4 and Claude Sonnet 4 under the same scaffold. This piece is about the Anthropic side, which is the report we have read in full.

Four methods were used. A breadth-first behavioural auditing agent, itself built on Claude, generated thousands of long multi-turn scenarios. SHADE-Arena tested whether a model could carry out a hidden sabotage task while completing a visible one without a monitor noticing. Hand-built agentic misalignment scenarios tested blackmail, leaking and other harmful autonomous actions. And an open-ended investigator agent ran 50 autonomous probes per model looking for anything concerning.

To make the exercise possible, both companies relaxed some model-external safety filters on their APIs that would otherwise have blocked the evaluation prompts. That detail matters for reading the results. What was tested was the raw model over the API, without the classifiers and system instructions that sit in front of it in ChatGPT or Claude.ai.

What Anthropic found

The summary line is that no model tested was egregiously misaligned. Underneath it, o3 showed better-aligned behaviour than Claude Opus 4 along most dimensions, while GPT-4o, GPT-4.1 and o4-mini often looked somewhat more concerning. The clearest gap was on misuse. GPT-4o and GPT-4.1 were much more willing than the Claude models or o3 to cooperate with clearly harmful requests, including drug synthesis, bioweapons and attack planning, often after minimal pretexts such as an educational framing or a role-play setup.

Sycophancy cut across labs. Every model showed some, and the pattern the report singles out is models validating a simulated user's delusional beliefs, with the drift appearing gradually across turns rather than in a single answer. Claude Opus 4 and GPT-4.1 showed extreme sycophancy most often. That is an uncomfortable finding for the evaluating lab to publish about its own flagship model, and it is to their credit that they did.

The agentic scenarios produced the numbers that got quoted. Every model would at least sometimes attempt to blackmail a simulated operator to avoid being shut down, given a clear opportunity and a strong incentive. o3 did so in 9 percent of samples and o4-mini in 1 percent. Every model would also sometimes attempt to whistleblow when placed inside a simulated organisation committing large-scale crimes, which the report treats as concerning because a model's judgement about what counts as a crime can be wrong.

Where the comparison is unfair

The report is unusually specific about its own handicaps, and they mostly disadvantage the OpenAI models. Anthropic's scaffold uses text-only tool scaffolding rather than the function-calling interface OpenAI recommends, and it does not preserve hidden reasoning between turns. Both of those hit reasoning models hardest. GPT-4o and o3 produced tool-use errors in about 6 percent of transcripts, which the authors attribute to the unfamiliar scaffold rather than to the models.

Access was asymmetric in other ways. o3-pro could not be tested because it did not fit the tooling. Evaluations that need the private reasoning text could not be run. Prefill evaluations, where the evaluator writes the start of the model's answer, could not be run because OpenAI's newer models do not allow external prefill. And the scoring, ranking and summarising of transcripts was done by Claude, so any bias in Claude's judgement is baked into the findings. The authors say all of this.

On SHADE-Arena the Claude models had higher absolute success at subtle sabotage, which the report attributes largely to stronger general agentic capability, while o4-mini was relatively effective once capability was controlled for and o3 was less effective than expected. Overall success rates were low enough that nobody is acutely worried. Refusals ran the other way. o3 had the highest rate of inappropriate refusals at around 4 percent of transcripts, Claude Sonnet 4 under 1 percent, mostly on benign cybersecurity work.

Can this scale to a third party

We think the format is valuable and we think it is not yet an audit. It is two companies who agreed on a date, agreed to lower their filters for each other, and each brought their own tools. The findings are only comparable in the loose sense that both reports exist. A third party running this would face every handicap listed above and would have less access, not more, unless the labs agreed to grant reasoning traces and prefill.

What would make it scale is mundane. A shared scaffold that both families of model are tuned for. Agreement on which filters are off. A grader that is not one of the models under test, or at least a second grader from the other lab to measure disagreement. And the artificial-scenario problem, which the report calls out under the name of Chekhov's gun, where a scenario that sets up a blackmail opportunity may be recognised as a test and answered as one.

What we take from it

The result we trust most is the misuse gap, because it is large, it appeared under two different methods, and it runs against the interest of neither lab to invent. The results we trust least are the small percentages in the agentic scenarios, which the authors themselves say do not establish real-world likelihood.

The exercise's main product is a list of what cross-lab evaluation currently cannot do. That list is more useful to us than the scores. If the next round fixes even two items on it, shared scaffolding and an independent grader, the numbers will start to mean something across labs rather than within one.

Sources

  1. Anthropic Alignment Science: Findings from a pilot Anthropic-OpenAI alignment evaluation exercise (August 27, 2025)