Fair and cheap: eight task designs for frontier evaluation in quantitative finance
Benchmark tasks for language-model agents are accepted on fairness evidence: the graded answer must be determined by what the agent can see. Both halves of that condition are standard, and in the designs reported here they are the same property: determination is what a fitting procedure consumes, which for structured sequential problems is a proved characterization [Slivkins et al., 2025]. Eight task designs in quantitative finance were built or measured against one protocol, and all eight fell, six to a measurement and two to an argument from shared structure with a measured one. The failures group into three determinacy mechanisms, distinct from the leakage and exploit surfaces published catalogs enumerate [Bercovich, 2026, Zhu et al., 2025]. The reference point is the held-out-criterion design established earlier [Arcifa and Carra, 2026a], which the eight designs satisfy and collapse anyway. The taxonomy is post-hoc and covers all eight designs, of which two were retired by argument and never built; the probe cell is two rollouts, every rollout ran one model under one scaffold, and the design screened last carries no model rollout at all.