Publications

Everything we find, published in full

Our preprints are released as they are finished, with the code, data, and evidence needed to examine them. Each carries a stable identifier so it can be cited before and after peer review. Reading one takes a quick sign-up; it tells us who's engaging with the work, nothing more.

MRF-2026-04

September 2026

Preprint Zenodo zenodo.22285450 First page of Fair and cheap: eight task designs for frontier evaluation in quantitative finance
680
views
139
downloads

Fair and cheap: eight task designs for frontier evaluation in quantitative finance

Dr. Ricardo Arcifa, Francieli Carra

Benchmark tasks for language-model agents are accepted on fairness evidence: the graded answer must be determined by what the agent can see. Both halves of that condition are standard, and in the designs reported here they are the same property: determination is what a fitting procedure consumes, which for structured sequential problems is a proved characterization [Slivkins et al., 2025]. Eight task designs in quantitative finance were built or measured against one protocol, and all eight fell, six to a measurement and two to an argument from shared structure with a measured one. The failures group into three determinacy mechanisms, distinct from the leakage and exploit surfaces published catalogs enumerate [Bercovich, 2026, Zhu et al., 2025]. The reference point is the held-out-criterion design established earlier [Arcifa and Carra, 2026a], which the eight designs satisfy and collapse anyway. The taxonomy is post-hoc and covers all eight designs, of which two were retired by argument and never built; the probe cell is two rollouts, every rollout ran one model under one scaffold, and the design screened last carries no model rollout at all.

MRF-2026-03

August 2026

Preprint Zenodo zenodo.22285449 First page of An Ability Label Raises the Effort of an Agent
1.1k
views
246
downloads

An Ability Label Raises the Effort of an Agent

Pre-Registered Expectancy Framing on a Held-Out-Criterion Task

Dr. Ricardo Arcifa, Francieli Carra

An ability label is a sentence that tells a worker how it is expected to perform before the work begins. In humans such labels move outcomes: a positive expectation raises performance (the Pygmalion effect) and a salient negative stereotype lowers it (stereotype threat). Whether a language-model agent responds to the same sentence, and in which direction, is an open question the human designs cannot answer, because they cannot hold ability constant across labeled groups. A model can: same weights, same task, same seeds, only the label differs. This five-stage, pre-registered study of one agentic task with a held-out grading criterion found no effect of group-form labels for Opus 5 or GPT-5.6 Sol at 20 seeds per cell. An individual-form negative label raised the reasoning tokens Opus 5 spent by about half a standard deviation (p = 0.0106 over 30 matched seeds): the negative label raised effort rather than depressing it, the opposite of the human prediction.

MRF-2026-02

July 2026

Preprint Zenodo zenodo.22285432 First page of Measuring Agent Self-Knowledge Under a Criterion Held Out of the Environment
1.5k
views
321
downloads

Measuring Agent Self-Knowledge Under a Criterion Held Out of the Environment

Dr. Ricardo Arcifa, Francieli Carra

Calibration results for tool-using agents largely measure access to verification rather than self-knowledge. We measure the complement: on three certified task families whose grading criterion is held out of the environment, the agent induces a hidden rule from labeled examples, is graded on cases the container never holds, and records an unrewarded declaration of its probability that the implementation generalizes exactly. Two frontier configurations ran every cell at 20 seeds. Both declare mean confidence 0.24 to 0.48 above their measured rate of exact generalization on two of the three families, and at most 0.13 above it on the third: overconfidence in this setting is a property of the family rather than a fixed trait of the model.

MRF-2026-01

June 2026

Preprint Zenodo zenodo.22285444 First page of ATLAS: Adversarial, Traceable, Latent-Criterion, Auditable, and Seed-Calibrated Task Acceptance
2.3k
views
534
downloads

ATLAS: Adversarial, Traceable, Latent-Criterion, Auditable, and Seed-Calibrated Task Acceptance

A Calibration Study for Agent Benchmark Task Families

Dr. Ricardo Arcifa, Francieli Carra

Benchmark tasks for language-model agents are typically accepted on the evidence that a reference solution passes the grader. This criterion is necessary and insufficient: it cannot detect graders that award reward for schema-conformant junk, copied inputs, forged reward files, or doing nothing. We describe an acceptance pipeline that treats task acceptance as an adversarial testing problem — static linting, an oracle baseline, a no-op baseline, a determinism check, a battery of scripted cheating agents, and a generalization check across seeds — with the outcome recorded in a portable certificate a benchmark consumer can inspect without trusting the author. We then calibrate three certified task families against two frontier laboratories at 20 seeds per cell.

Related work is welcome.

We welcome replications, critiques, and collaborations on any of the work above.

Get in touch