Publications MRF-2026-01

Preprint MRF-2026-01 33 pages Zenodo DOI 10.5281/zenodo.22285444

ATLAS: Adversarial, Traceable, Latent-Criterion, Auditable, and Seed-Calibrated Task Acceptance

A Calibration Study for Agent Benchmark Task Families

Dr. Ricardo Arcifa, Francieli Carra · Montana Research Foundation

Read the paper (PDF) GitHub Code and data
2.3k
views
534
downloads

Abstract

Benchmark tasks for language-model agents are typically accepted on the evidence that a reference solution passes the grader. This criterion is necessary and insufficient: it cannot detect graders that award reward for schema-conformant junk, copied inputs, forged reward files, or doing nothing. We describe an acceptance pipeline that treats task acceptance as an adversarial testing problem — static linting, an oracle baseline, a no-op baseline, a determinism check, a battery of scripted cheating agents, and a generalization check across seeds — with the outcome recorded in a portable certificate a benchmark consumer can inspect without trusting the author. We then calibrate three certified task families against two frontier laboratories at 20 seeds per cell.

Licensed under CC BY 4.0. Every number in the paper traces to a committed artifact named in the text. The run records and a verify.py that recomputes the headline figures are in the open-data repository.

Cite this preprint

Arcifa, R., & Carra, F. (2026). ATLAS: Adversarial, Traceable, Latent-Criterion, Auditable, and Seed-Calibrated Task Acceptance. Montana Research Foundation preprint MRF-2026-01. https://doi.org/10.5281/zenodo.22285444

BibTeX
@techreport{arcifa2026atlas,
  title = {ATLAS: Adversarial, Traceable, Latent-Criterion, Auditable, and Seed-Calibrated Task Acceptance},
  author = {Arcifa, Ricardo and Carra, Francieli},
  institution = {Montana Research Foundation},
  type = {Preprint},
  number = {MRF-2026-01},
  year = {2026},
  month = {6},
  doi = {10.5281/zenodo.22285444},
  url = {https://montanaresearch.org/publications/mrf-2026-01/},
  note = {PDF: https://montanaresearch.org/download/paper/mrf-2026-01}
}