Research

Independent AI and ML research, openly shared

We take on the questions we think matter most, and give them the time they deserve. Across all four areas below, that means evaluation: building the methods and benchmarks that tell us whether an AI system's claims hold up. Our work is published openly so anyone can examine it, challenge it, and build on it.

Research areas

Four areas, one standard

Each area feeds the same thing: evaluation methodology that tells us whether an AI system's claims hold up.

Foundations of AI

Fundamental questions about how modern AI systems learn, generalize, and fail. The answers are what any evaluation has to be built on.

Trustworthy Systems

Methods for evaluating whether an AI system's behaviour can be understood, measured, and relied upon.

Applied Research

Turning evaluation methodology into benchmarks and task designs that hold up under real-world constraints.

Open Science

Publishing the evaluation data, code, and methodology behind our findings so other labs can verify and reuse them.

Every preprint we publish ships with its run records and a verification script in our public open-data repository.

Task benchmarking

How we build a task benchmark

A task benchmark is a small adversarial system: an environment, an instruction, and a grader that has to withstand an agent trying to satisfy it for the wrong reasons. We don't accept a task because a reference solution passes it. Every task goes through six checks built to catch a grader that can be gamed, then gets calibrated against real models to see whether it actually tells them apart.

The six checks

Static lint
Oracle baseline
No-op baseline
Determinism check
Cheating-agent battery
Generalization check
✓ Certificate

Calibrated against two models

Three certified task families, twenty runs each, no stopping rule — the tasks separate the models by different amounts, and one doesn't separate them at all:

rule-induction Claude (Opus 5) 5 of 20
rule-induction GPT (GPT-5.6 Sol) 1 of 20
format-induction Claude (Opus 5) 12 of 20
format-induction GPT (GPT-5.6 Sol) 3 of 20
protocol-induction Claude (Opus 5) 11 of 20
protocol-induction GPT (GPT-5.6 Sol) 10 of 20

Passed Failed

Every graded criterion

Each family grades one deciding criterion plus three partial ones that name how a failure fell short — the exact per-seed value of all four, from the same run records the paper's own tables and figures are built from. Failures aren't failed broadly: a criterion is usually satisfied even on a seed that fails overall, and the ones that dim below name where that specific seed actually fell short.

rule-induction Claude (Opus 5)

exact_rule
boundary_agreement
interaction_agreement
tier_floor

rule-induction GPT (GPT-5.6 Sol)

exact_rule
boundary_agreement
interaction_agreement
tier_floor

format-induction Claude (Opus 5)

exact_spec
positive_agreement
negative_agreement
constraint_floor

format-induction GPT (GPT-5.6 Sol)

exact_spec
positive_agreement
negative_agreement
constraint_floor

protocol-induction Claude (Opus 5)

exact_protocol
accepted_agreement
rejected_agreement
rule_floor

protocol-induction GPT (GPT-5.6 Sol)

exact_protocol
accepted_agreement
rejected_agreement
rule_floor

Pass Partial Fail

Five early rule-induction / Claude seeds were graded under an earlier rubric that didn't record the three partial criteria. Seed 0 passed overall, so those criteria are filled in as pass — every criterion is 1 on a passing seed, by construction. Seeds 1-4 failed overall with the partial criteria genuinely unrecorded, so those four columns are left blank in those rows rather than assigned an outcome that was never measured.

Read the full calibration study in ATLAS

By science

Different fields, the same discipline

Evaluation looks different in every field, but the standard doesn't move: a claim needs to survive an attempt to break it before it counts as a result.

Mathematics

The modeling and proof techniques behind rigorous evaluation design: what a task's ground truth actually depends on, and how to state it precisely enough to grade.

Physics

Simulation and dynamical systems: domains where an answer has to hold up against physical law, not just a held-out test set.

Chemistry

Molecular structure and property prediction, where a wrong answer costs a real synthesis attempt to find out about.

Biology & Medicine

From protein structure to clinical reasoning, fields where the bar for a correct answer is set by consequences.

Computer Science

Systems, security, and software engineering — the domain most of our published evaluation work comes from so far.

Earth & Environmental Science

Planetary and environmental data, where ground truth is expensive to collect and easy to get quietly wrong.

Datasets

In progress: synthetic data for fields real data can't easily reach

We're building a synthetic dataset for cancer cell-line perturbation research, decomposed into verifiable generative parts and checked against real data at every stage. It isn't published yet. Synthetic datasets for financial markets and neuroscience are next.

You're welcome to contribute.

We are always interested in hearing from researchers who share our approach.

Get in touch