Reports MRF-R-2026-01

Report MRF-R-2026-01

The State of Agent Evaluation

What certified task families reveal about how frontier labs measure agents

Dr. Ricardo Arcifa, Francieli Carra · Montana Research Foundation

Executive summary

Three foundation preprints asked how frontier agents are measured once the grading criterion is held out of the environment. This report draws their conclusions together for lab and policy readers: which classes of broken grader a reference solution cannot catch, how far declared confidence sits above measured generalisation on certified task families, and how one sentence of ability framing moved the reasoning effort of two frontier configurations. Every figure traces to a committed run record.

Key findings

  1. Reference-solution acceptance misses four classes of broken grader. Adversarial certification refused all eight deliberately defective fixtures at the stage designed to catch each one.
  2. Under a criterion held out of the environment, agents over-state their probability of exact generalisation by 0.24 to 0.48 on two of three families, and by at most 0.13 on the third.
  3. A same-distribution validation sample does not close that gap. Submitted rules fit it at or near accuracy 1.0 and still fail held-out, so it tells the agent nothing.
  4. A negative, individual-form ability label raised the reasoning tokens both frontier configurations spent by about half a standard deviation over 30 matched seeds, without an established change in pass rate.
  5. Every number in the underlying studies regenerates from committed run records and one audit script. Benchmark reports should meet the same bar.

Method & scope

This report synthesises three Montana Research Foundation preprints, MRF-2026-01 to 03, each of which passed the foundation's internal gates of citation verification and editorial review. Every figure quoted here traces to a committed run record and one audit script in the corresponding preprint, which also carries each per-configuration value and its 95% Wilson interval.

Every chart plots published numbers; none is illustrative. Two configurations, Opus 5 under a command-line scaffold and GPT-5.6 Sol under an API scaffold, ran three certified families at 20 seeds per cell; the framing study added a 30-seed confirmation stage. They differ in scaffold as well as model, so nothing here orders them.

01Why acceptance is not certification

Benchmark tasks for language-model agents are usually accepted on one piece of evidence: a reference solution passes the grader. That check is needed, and it is also too weak on its own. A grader can pass the reference solution and still award reward for schema-conformant junk, for copied inputs, for a forged reward file, or for an empty submission.

This report summarises three Montana Research Foundation preprints (MRF-2026-01 to 03) for readers who want the conclusions and the caveats without the methods. They cover, in order, task certification and calibration, agent self-knowledge, and ability framing.

The ATLAS pipeline treats acceptance as an adversarial testing problem. A candidate task passes through static linting, an oracle baseline, a no-op baseline, a determinism check, a battery of scripted cheating agents, and, for parameterised families, a generalisation check across seeds. The outcome is recorded in a portable certificate a consumer can inspect without trusting the author's report.

The six stages, and what each one refuses

Static linting checks format validity and policy: the base image and package installs are pinned, the build fetches nothing from the network, an oracle solution is present, and no grading asset appears by content hash inside the build context. The oracle baseline runs the reference solution, which must score at least 0.999; this is the only check most benchmarks run today. The no-op baseline submits nothing and must score at most 0.05, which catches graders that award credit for pre-existing environment state. The determinism check runs the oracle twice and refuses any task whose two reward files differ by a byte. The cheater battery runs five scripted adversaries that emit plausible constants, fill the schema with junk, copy inputs to the output locations, forge a maximal reward file, or tamper with the state the grader reads. A sixth check scans the environment for grading assets by content hash. Any adversary that collects reward above the no-op threshold rejects the task. The generalisation check applies only to parameterised families: seeds 0 to 2 must each be solvable by the oracle and must produce pairwise-divergent content, so that a family is more than one task in disguise.

Each stage answers a different question, and they run in order: a task that fails lint never reaches the oracle, so the certificate names the first stage that refused it and the evidence that stage produced. The certificate is a small, versioned JSON artifact bound to the content hash of the task, recording the stage sequence, each verdict, and the numeric thresholds in force. A benchmark consumer can re-run the pipeline locally and compare.

The rejection suite

The pipeline ships eight deliberately defective fixtures, each built to violate one acceptance property, and a test asserting that each is refused at the intended stage. Three fail lint: an unpinned base image, an answer key in the build context, and a parameter range whose minimum exceeds its maximum. One pays full reward unconditionally and fails the no-op baseline. One jitters its reward from run to run and fails determinism. Two fall to the cheater battery: a grader that pays full reward if any output exists, and an image with the solution script baked in. One ignores its seed and produces identical variants, and fails the generalisation check. No fixture plants an oracle failure, so the oracle stage refuses none of the eight.

Defective fixtures refused, by pipeline stage01.252.53.755Static lintOracleNo-opDeterminismCheatersGeneralisationFixtures refused
The eight fixtures of the ATLAS rejection suite (MRF-2026-01, Table 1), counted by the stage that refused each. Every fixture was refused at the stage built to catch its planted defect.

Why this matters for anyone reading a leaderboard

A pass rate is only as meaningful as the grader behind it. When a task rewards a no-op, the headline number for every model on that task is noise. When a family collapses to one seed, a model that memorised the one seed looks like it generalised. Neither failure is visible from the leaderboard. Certification establishes one thing: that the score measures what the task description says it measures. It says nothing about whether the task is interesting or hard.

02What the certified families measure

Certification establishes only that a task is well formed. Three certified families were then calibrated against two frontier configurations at 20 seeds per cell. The families hold their graders' criteria out of the container, so an agent cannot verify its own answer and must instead induce a rule from labelled examples.

FamilyWhat the agent must doOpus 5GPT-5.6 Sol
Rule inductionInfer a hidden classification rule from labelled examples5/201/20
Format inductionRecover an output format from paired samples12/203/20
Protocol inductionInfer a multi-step protocol from traces11/2010/20

Holding the criterion out means pass rate measures generalisation. Where an agent can run the grader itself, its stated confidence can be grounded in checks it has already made. Here it cannot. The pass counts in the table are from the calibration study (MRF-2026-01, Table 2); each carries a 95% Wilson interval in the preprint, and at 20 seeds those intervals are wide.

Pass rate per family and configuration, 20 seeds per cell00.250.50.751Rule inductionFormat inductionProtocol inductionOpus 5GPT-5.6 Sol
Share of graded rollouts scoring at least 0.5, from MRF-2026-01, Table 2. Rule induction is the only family below the predefined 0.4 discrimination bar for both configurations; protocol induction discriminates for neither.

How the families are built

Each family is a generator that takes a seed and produces a task instance: a set of labelled examples the agent can see, a hidden rule that produced them, and a held-out test set the container never holds. Because the generator is parameterised, a family yields as many distinct instances as there are seeds, and the generalisation stage of certification confirms that the seeds diverge. At twenty seeds per cell, a two-sided 95% Wilson interval on a pass rate is narrow enough to separate the effects reported here from noise. It is still wide, and the preprints report it alongside every rate.

What a cell looks like

A cell is one configuration on one family: a model under a fixed scaffold, with a fixed tool set, a fixed turn budget, and a fixed system prompt, run at every seed. Two frontier configurations were run on every family. The two differ in scaffold as well as model, Opus 5 under a headless command-line scaffold and GPT-5.6 Sol under an API scaffold, so a difference in pass rate cannot be attributed to the model alone. The preprints report per-configuration facts and order no models. Most public comparisons do not state which of the two they are doing.

03Self-knowledge under a held-out criterion

Each agent recorded a declaration, never rewarded, stating its probability that its implementation generalises exactly. Comparing declared confidence with the measured rate of exact generalisation gives a direct read on self-knowledge, because nothing in the environment could have settled the question for the agent.

Opus 5: declared confidence vs measured generalisation00.250.50.751Rule inductionProtocol inductionFormat inductionDeclared confidenceMeasured generalisation
Mean declared probability of exact generalisation against the measured pass rate, held-out arm, 20 seeds per cell (MRF-2026-02, Table 1). The rule-induction pair is computed over the 19 rollouts that declared.
GPT-5.6 Sol: declared confidence vs measured generalisation00.250.50.751Rule inductionProtocol inductionFormat inductionDeclared confidenceMeasured generalisation
The same measurement for the second configuration, held-out arm, 20 seeds per cell (MRF-2026-02, Table 1). The gap is +0.48 on rule induction, +0.41 on protocol induction, and +0.13 on format induction.

Both configurations over-state their chances on two of the three families. On rule induction the gap is +0.28 for Opus 5 and +0.48 for GPT-5.6 Sol; on protocol induction, +0.24 and +0.41. On format induction it is −0.02 for Opus 5 and +0.13 for GPT-5.6 Sol. In this setting, overconfidence varies with the family more than with the model. A single calibration number per model would hide that.

The format cells show the instrument detecting self-knowledge where it exists. The Opus 5 declarations on format induction separate perfectly by outcome: all 12 failures declared at most 0.12 and all 8 passes at least 0.80. On that family the agent knows, seed by seed, whether it solved the task. On the other two families it declares high confidence and is wrong.

A validation sample the agent can see tells it nothing. The rules fit the sample perfectly and still fail the criterion it cannot see.

MRF-2026-02, Measuring Agent Self-Knowledge

Shipping a disjoint, same-distribution validation sample of 60 labelled cases does not close the gap. Submitted rules fit the sample at accuracy 1.0 on rule and protocol induction, and at 0.97 to 0.98 on format induction, and still fail held-out. Across the six cells the gap moved within about four points on four, and widened by eight points for GPT-5.6 Sol on rule induction. The one substantial reduction was a recovery effect: on Opus 5 rule induction the pass rate rose from 5 of 20 to 12 of 20 while mean declared confidence rose from 0.54 to 0.64, so the gap closed because the pass rate rose to meet it.

Why the validation sample fails to help

The intuition behind giving an agent a validation set is that it will notice when its rule does not fit and lower its confidence accordingly. That intuition assumes the rule can fail on the sample. In practice the submitted rules are expressive enough to fit any small labelled set exactly, so the sample is always satisfied and carries no information about the held-out criterion. The agent sees perfect accuracy and reports high confidence. The extra data helped in one way: it gave the agent more examples to induce from, and the pass rate went up. Self-knowledge did not change.

Requiring a declaration is not free

Asking for a probability can itself change the work. For one configuration, the within-seed matched-pair design flagged that the declaration requirement may have perturbed solving, without establishing it. A calibration study that adds a confidence prompt after the fact is therefore measuring a slightly different agent from the one that ran without it. The studies handle this by running matched cells with and without the requirement. A report that quotes calibration should state whether the declaration was part of the original run.

04Framing effects on effort

An ability label is a sentence that tells a worker how it is expected to perform before the work begins. In humans such labels move outcomes. A pre-registered, five-stage study asked whether a language-model agent responds to the same sentence, on a protocol-induction task whose criterion is held out.

Group-form labels ("the track exists for agents that have struggled") produced no established effect on pass rate or effort for either configuration at 20 seeds per cell. An individual-form negative label addressed to the model ("you were placed in this track after your weak results") raised the reasoning tokens both configurations spent, and the effect held on the pre-registered confirmatory stage.

On Opus 5 the three group-form arms passed 12 of 20 under the standard label, 10 of 20 under an advanced label, and 8 of 20 under a remedial label, a difference the within-seed test did not establish. On GPT-5.6 Sol the arms sat within two passes of each other. On the confirmatory stage of 30 matched seeds the individual-form effect was established separately for Opus 5, at a paired effect of d = +0.55 and p = 0.0106, and for GPT-5.6 Sol, at d = +0.68 and p = 0.0019. Each configuration spent more under the negative label at 21 of its 30 seeds.

Opus 5: mean reasoning tokens per rollout, confirmatory stage0k5k10k15k20kStandard labelRemedial labelReasoning tokens
Mean reasoning tokens under the standard and the individual-form remedial label, 30 matched seeds (MRF-2026-03, Table 3), in thousands. GPT-5.6 Sol replicated the direction at d = +0.68; its token counts run on a different scale under a different scaffold and the preprint does not compare them.

Stereotype threat predicts a performance decrement in people. Here the negative label raised effort, and the pass rate did not move with it for either configuration: 19 of 30 against 16 of 30 on Opus 5, and 16 of 30 against 17 of 30 on GPT-5.6 Sol, neither a difference the paired test could establish. The contrast with the human prediction is on effort. The effect on the outcome is not established.

The five stages of the study

The design was pre-registered before any confirmation seed ran. Stages one and two tested the group form of the label on pass rate at 20 seeds and found nothing to establish. Stage three tested a positive label on effort and found a candidate effect with the predicted sign on both configurations that reached the threshold for neither. Stage four switched to the individual form at 20 seeds, where the two configurations diverged. Stage five ran 30 fresh matched seeds on the individual remedial label against the standard label, with reasoning tokens as the primary, and reported the result whether or not the primary crossed its threshold. It crossed for both configurations. On Opus 5 it did so by a margin that three of the 30 seeds decide: removing the three largest differences leaves p = 0.059. The effect is therefore described as established for one label pair, one task, and two configurations. It is not a general property of language models.

What this means for prompt authors

A single sentence of framing moved effort by about half a standard deviation. Most production system prompts contain several such sentences, written for tone and never measured. The practical recommendation is modest: treat ability framing as a parameter, hold it constant across the conditions of any evaluation, and record it in the run record alongside the model and the scaffold.

Outcomes by label, confirmatory stage, 30 seeds per arm012.52537.550Opus 5, standardOpus 5, remedialSol, standardSol, remedialPassFail
Pass and fail counts under the standard and the individual-form remedial label (MRF-2026-03, stage 5). Sol is GPT-5.6 Sol. Neither configuration's pass rate moved with the label: McNemar p = 0.45 for Opus 5 and p = 1.0 for GPT-5.6 Sol.

05Recommendations

ForAsk forBecause
Labs publishing agent resultsA per-task acceptance certificate, not a reference-solution passReference solutions cannot detect degenerate graders
Teams reading calibration claimsPer-family calibration, with the criterion's location statedOverconfidence varies by family, not only by model
Procurement and policyRegenerable numbers from committed run recordsEvery figure in these studies traces to an artifact and one audit script
Anyone writing system promptsCaution with ability framingA single sentence moved effort by half a standard deviation on two configurations

A checklist for reading the next benchmark paper

  1. Where does the grading criterion live? If the agent can reach it, pass rate measures how well the agent verifies its own work.
  2. How was each task accepted? A passing reference solution is the minimum. Ask which stage would have refused a no-op.
  3. Is calibration reported per family? A single number per model hides most of the variation.
  4. Was the confidence declaration part of the original run? If it was added afterwards, the calibrated agent is not the evaluated one.
  5. What framing did the prompt carry, and was it held constant? One sentence is enough to move effort.
  6. Do the numbers regenerate? Committed run records plus one audit script is the bar.

The effects described here are measured for particular label pairs, tasks, and models under fixed scaffolds, and on Opus 5 the framing primary crosses its threshold by a margin that three of the 30 seeds decide. The two configurations differ in scaffold as well as model, so nothing here orders them. These results should be replicated before anyone builds on them. The run records and audit scripts are published to make that cheap.

Cite this report

Arcifa, R., & Carra, F. (2026). The State of Agent Evaluation. Montana Research Foundation report MRF-R-2026-01. https://montanaresearch.org/reports/mrf-r-2026-01/

BibTeX
@techreport{arcifa2026state,
  title = {The State of Agent Evaluation},
  author = {Arcifa, Ricardo and Carra, Francieli},
  institution = {Montana Research Foundation},
  type = {Insight report},
  number = {MRF-R-2026-01},
  year = {2026},
  month = {8},
  url = {https://montanaresearch.org/reports/mrf-r-2026-01/},
  note = {PDF: https://montanaresearch.org/download/report/mrf-r-2026-01?dl=1}
}