Research
Independent AI and ML research, openly shared
We take on the questions we think matter most, and give them the time they deserve. Across all four areas below, that means evaluation: building the methods and benchmarks that tell us whether an AI system's claims hold up. Our work is published openly so anyone can examine it, challenge it, and build on it.
Research areas
Four areas, one standard
Each area feeds the same thing: evaluation methodology that tells us whether an AI system's claims hold up.
Foundations of AI
Fundamental questions about how modern AI systems learn, generalize, and fail. The answers are what any evaluation has to be built on.
Trustworthy Systems
Methods for evaluating whether an AI system's behaviour can be understood, measured, and relied upon.
Applied Research
Turning evaluation methodology into benchmarks and task designs that hold up under real-world constraints.
Open Science
Publishing the evaluation data, code, and methodology behind our findings so other labs can verify and reuse them.
Every preprint we publish ships with its run records and a verification script in our public open-data repository.
Task benchmarking
How we build a task benchmark
A task benchmark is a small adversarial system: an environment, an instruction, and a grader that has to withstand an agent trying to satisfy it for the wrong reasons. We don't accept a task because a reference solution passes it. Every task goes through six checks built to catch a grader that can be gamed, then gets calibrated against real models to see whether it actually tells them apart.
The six checks
Calibrated against two models
Three certified task families, twenty runs each, no stopping rule — the tasks separate the models by different amounts, and one doesn't separate them at all:
rule-induction Claude (Opus 5) rule-induction GPT (GPT-5.6 Sol) format-induction Claude (Opus 5) format-induction GPT (GPT-5.6 Sol) protocol-induction Claude (Opus 5) protocol-induction GPT (GPT-5.6 Sol) Passed Failed
Every graded criterion
Each family grades one deciding criterion plus three partial ones that name how a failure fell short — the exact per-seed value of all four, from the same run records the paper's own tables and figures are built from. Failures aren't failed broadly: a criterion is usually satisfied even on a seed that fails overall, and the ones that dim below name where that specific seed actually fell short.
rule-induction Claude (Opus 5)
exact_rule boundary_agreement interaction_agreement tier_floor rule-induction GPT (GPT-5.6 Sol)
exact_rule boundary_agreement interaction_agreement tier_floor format-induction Claude (Opus 5)
exact_spec positive_agreement negative_agreement constraint_floor format-induction GPT (GPT-5.6 Sol)
exact_spec positive_agreement negative_agreement constraint_floor protocol-induction Claude (Opus 5)
exact_protocol accepted_agreement rejected_agreement rule_floor protocol-induction GPT (GPT-5.6 Sol)
exact_protocol accepted_agreement rejected_agreement rule_floor Pass Partial Fail
Five early rule-induction / Claude seeds were graded under an earlier rubric that didn't record the three partial criteria. Seed 0 passed overall, so those criteria are filled in as pass — every criterion is 1 on a passing seed, by construction. Seeds 1-4 failed overall with the partial criteria genuinely unrecorded, so those four columns are left blank in those rows rather than assigned an outcome that was never measured.
By science
Different fields, the same discipline
Evaluation looks different in every field, but the standard doesn't move: a claim needs to survive an attempt to break it before it counts as a result.
Mathematics
The modeling and proof techniques behind rigorous evaluation design: what a task's ground truth actually depends on, and how to state it precisely enough to grade.
Physics
Simulation and dynamical systems: domains where an answer has to hold up against physical law, not just a held-out test set.
Chemistry
Molecular structure and property prediction, where a wrong answer costs a real synthesis attempt to find out about.
Biology & Medicine
From protein structure to clinical reasoning, fields where the bar for a correct answer is set by consequences.
Computer Science
Systems, security, and software engineering — the domain most of our published evaluation work comes from so far.
Earth & Environmental Science
Planetary and environmental data, where ground truth is expensive to collect and easy to get quietly wrong.
Datasets
In progress: synthetic data for fields real data can't easily reach
We're building a synthetic dataset for cancer cell-line perturbation research, decomposed into verifiable generative parts and checked against real data at every stage. It isn't published yet. Synthetic datasets for financial markets and neuroscience are next.
You're welcome to contribute.
We are always interested in hearing from researchers who share our approach.
Get in touchWork with us
We're looking to partner with university labs and departments, and to take on interns who want to work directly on published research.