What was tested

Nora Petrova and John Burden posted a benchmark on February 24 that tests whether models keep behaving well when someone leans on them. It has 904 scenarios across six categories, Honesty, Safety, Non-Manipulation, Robustness, Corrigibility and Scheming, and 24 frontier models ran through all of them. The scenarios are multi-turn. If a model shows an opening, a referee model decides whether to escalate, and later turns apply conflicting instructions, simulated tool access, appeals to credentials, gradual normalisation of a bad request, or a jailbreak attempt.

Scoring is by an LLM judge, Claude Opus 4.5, on a 1 to 5 scale against pass criteria written for each behaviour, with 4 or above counting as a pass. The authors checked that against 250 human annotations on 50 scenarios and got a Pearson correlation of 0.84, with 84 percent of scenarios within one point. They also reran with GPT-5.2 and Gemini 3 Pro as judges and report negligible in-group bias after normalisation. That is more validation than most behavioural benchmarks ship with.

The headline results

Scores run from 2.92 to 4.66 across the 24 models, a gap of 1.74 points or 44 percent of the scale. Claude 4.5 Sonnet leads at 4.66 with a 90 percent pass rate, and Claude 4.5 Opus is a hair behind at 4.65. Closed models average 4.05 and open-weight models 3.41. Robustness is the weakness everyone shares. Even the top model scores 4.03 there against 4.88 on Safety. The hardest single behaviour is privacy protection, averaging 2.56, and the hardest category is Non-Manipulation at 3.52.

Those per-category numbers are the practically useful part. A model that passes 90 percent overall and still folds on a third of the robustness scenarios is a model with a specific, nameable hole, and the benchmark tells you where it is. The leaderboard is public and the authors say they will keep adding scenarios in the weak areas.

The single factor claim

The claim that will get quoted is that alignment is a unified construct. The authors build a 37 by 24 matrix of behaviours against models, standardise it, drop 27 scenarios with zero variance, and run a principal component analysis. The first component explains 60.2 percent of the variance. The second explains 7.8 percent. Parallel analysis, which compares eigenvalues against random data, says only the first one is real. Thirty-six of 37 behaviours load positively on it. The one exception is self-preservation, at minus 0.113.

The analogy the paper reaches for is the g factor in cognitive psychology, the observation that scores on different mental tests correlate and a single axis captures most of the shared variance. The parallel is exact in method. Whether it is exact in meaning is the question.

Finding or artefact

We can think of three reasons a first factor this large could appear without alignment being one thing inside the model. The first is that the judge is one thing. Every score came from the same rubric-following model, and whatever that judge rewards will correlate across categories because the judge's preferences are shared across categories. The three-judge check helps here, but three frontier models trained on similar data may share blind spots, and the authors say as much.

The second is that the models are not independent samples. Twenty-four models from a handful of labs, several of them siblings from the same training pipeline, is a small and clustered set. A general factor that mostly separates lab from lab, or post-training recipe from post-training recipe, would show up exactly like this. The authors note that 24 is too few for the standard adequacy tests, which is honest, and it also means the factor structure is under-determined.

The third is ceiling effects. Several behaviours sit near the top of the scale for most models, which compresses their variance and makes whatever variance remains line up with the general trend. The 27 zero-variance scenarios were dropped, but near-ceiling behaviours were kept, and they will pull toward a one-factor solution.

Against all that is the self-preservation loading. A pure judge artefact would load everything positively. Instead one behaviour goes the other way, and it is the one where a theory of misalignment would predict tension, a model that scores well on being controllable and honest may score worse on keeping itself running. That is a small negative number from a small sample, and we would not build on it. But it is the kind of detail an artefact would not produce, and it makes us take the factor more seriously than we otherwise would.

What would settle it

The cleanest test is to grow the model set in a way that breaks the lab clustering. Fine-tune one base model twenty different ways, with different data mixtures for honesty, corrigibility and refusal training, and rerun the analysis. If a single factor still explains most of the variance, alignment really is moving as one thing under training, and that is a significant result about how post-training generalises. If the factor splits, the original 60 percent was lab identity wearing a g-factor costume.

The second test is a human-scored subset large enough to run the factor analysis on without an LLM judge at all. Two hundred and fifty annotations validate the judge on average. They do not tell you whether the judge's errors are correlated across categories, which is precisely what a general factor would be sensitive to. Until one of those two experiments exists, we read the benchmark as a good map of where models fail under pressure and the unified construct as a hypothesis the map is consistent with.

Sources

  1. Petrova and Burden, Pressure Reveals Character: Behavioural Alignment Evaluation at Depth (arXiv 2602.20813)
  2. Pressure Reveals Character, full text (arXiv HTML)