Reports

Findings, written for decision-makers

Our reports distil what we learn into something a lab lead, a policy team, or a board can act on. Each one cites the studies behind every number. Reading or downloading one takes a quick sign-up, so we know who's acting on the findings.

MRF-R-2026-03

August 2026

Report First page of Agents in Financial Services
960
views
188
downloads

Agents in Financial Services

Adoption, customer-facing risk, model error rates, and the rules already in force

Dr. Ricardo Arcifa, Francieli Carra

Three quarters of UK financial firms already use AI, a third say they fully understand the systems they run, and more than a third of the US population has dealt with a bank's chatbot. Against that, a tribunal has held an airline liable for its chatbot's advice, a finance question-answering benchmark found a frontier model wrong or silent on four questions in five, and the EU has classified credit scoring as high-risk. This report puts the numbers together and sets out controls a risk function can sign off.

MRF-R-2026-02

August 2026

Report First page of Coding Agents in Production
1.3k
views
297
downloads

Coding Agents in Production

What controlled trials, delivery data, and the benchmarks themselves say about AI-assisted software engineering

Dr. Ricardo Arcifa, Francieli Carra

Engineering leaders are deciding how far to lean on coding agents on the strength of vendor demos and leaderboard scores. The published evidence is more mixed than either suggests: two controlled trials point in opposite directions, delivery telemetry shows stability falling as adoption rises, and the benchmark most often quoted has been retired by one of its own maintainers. This report sets out that evidence and the measurements a team should run on its own codebase before it changes how it works.

MRF-R-2026-01

August 2026

Report First page of The State of Agent Evaluation
1.8k
views
412
downloads

The State of Agent Evaluation

What certified task families reveal about how frontier labs measure agents

Dr. Ricardo Arcifa, Francieli Carra

Three foundation preprints asked how frontier agents are measured once the grading criterion is held out of the environment. This report draws their conclusions together for lab and policy readers: which classes of broken grader a reference solution cannot catch, how far declared confidence sits above measured generalisation on certified task families, and how one sentence of ability framing moved the reasoning effort of two frontier configurations. Every figure traces to a committed run record.

Briefings are available.

We can walk your team through any report, or scope one for the questions you're facing.

Get in touch