One number to rule them all: the Epoch Capabilities Index
Epoch has stitched nearly forty benchmarks into a single scale using an item response model, with Claude 3.5 Sonnet fixed at 130 and GPT-5 at 150. A method note on what that scale can compare, what it cannot, and the risks of a headline capability number.
The problem it is solving
Benchmarks saturate. GSM8K, HellaSwag, WinoGrande and MMLU were all informative once and now sit near their ceilings for any frontier model, while the benchmarks that can still separate models did not exist when the older models were evaluated. So there is no single test that covers both GPT-3.5 and the current frontier, and any attempt to draw a progress curve across years has to splice results from different tests with different scales.
Epoch's answer, published on October 28, is the Epoch Capabilities Index. It takes nearly forty benchmarks, fits a statistical model to every model-by-benchmark result Epoch has collected, and outputs one capability number per model. The scale is anchored so that Claude 3.5 Sonnet scores 130 and GPT-5 at medium reasoning scores 150. Models need at least four benchmark results to be included, and models from before 2023 are excluded because their data is too sparse to estimate reliably, which Epoch says is a data limitation rather than a technical one.
How item response theory gets applied to models
The model comes from educational testing. In item response theory a test-taker has a latent ability, each item has a difficulty and a discrimination, and the probability of a correct answer is a sigmoid of discrimination times ability minus difficulty. Epoch treats each AI model as a test-taker and each benchmark as an item, so the fitted equation in their public repository is performance equals sigmoid of discriminability times capability minus difficulty. Three sets of parameters come out. A capability score per model, which is the ECI. A difficulty per benchmark, which they call the EDI. And a discriminability per benchmark, which is the slope of that benchmark's curve.
The trick that makes this work across saturation is overlap. A model evaluated on both MMLU and a newer benchmark links the two scales, and a chain of such overlaps links the oldest benchmarks to the newest. Difficulty is inferred from which models pass rather than asserted in advance. Uncertainty comes from a bootstrap that resamples each model's results with replacement, refits, and re-anchors every draw so the two reference models sit at exactly 130 and 150. The confidence interval is then quantiles over draws. The code is public under an MIT licence and the underlying method is written up in a paper titled A Rosetta Stone for AI Benchmarks.
A worked example of what the fit means. If a benchmark has high discriminability, models just above its difficulty score well and models just below score badly, so it is a sharp instrument at one point on the scale and useless elsewhere. A benchmark with low discriminability spreads its information thinly across the whole range. The ECI weights each result by how much it tells you, which is why a perfect score on a saturated test contributes almost nothing and a partial score on a hard test contributes a lot.
What the number can and cannot compare
Epoch is candid about the limits, and we would rather repeat their caveats than invent softer ones. The absolute values are meaningless on their own, and the scale is linear and arbitrary, so 150 is not fifteen percent better than 130 in any sense you could cash out. Values are not linearly related to benchmark accuracy. A model specialised for one domain can score low despite being the best available at its task, because the index estimates broad capability and reads a weak showing elsewhere as low ability. And there is no interpretation of the form a score of X means the model can do Y, which Epoch explicitly declines to offer.
Two data problems are worth more attention than they got. The index relies partly on developer-reported scores, which Epoch acknowledges may be cherry-picked to flatter a model, and a developer who optimises for a benchmark in the index inflates its own estimate. The public post also notes that open-weight models appear about an order of magnitude behind closed models in training compute at a given ECI, and suggests this might be an artefact of spotty coverage of smaller models rather than a real gap. When the index's own authors are unsure whether a visible pattern is signal or sampling, the pattern should be reported with that doubt attached.
The domain sub-indices, for maths, software engineering and cyber, are scaled identically to the general index, which Epoch says prevents comparing long-run progress trends across domains. That is a real loss, because the question people most want the index for is whether coding is moving faster than maths, and the current design cannot answer it.
The risk of a headline number
The index is useful for what it was built for, which is a progress curve that survives benchmark turnover. It is less useful for the thing it will be used for, which is a leaderboard. A single number invites ranking, ranking invites optimisation, and a model trained to raise its ECI would do so by targeting the benchmarks with highest discriminability near its current score, which is exactly the contamination pattern the index's own caveats worry about. The bootstrap intervals are the defence, and we would ask anyone quoting an ECI to quote the interval with it.
The deeper concern is construct validity. IRT assumes a single latent trait, and the fit will produce one whether or not the trait exists. An index that combines maths, coding, knowledge and agentic tasks into one axis has decided that those are one thing. That may be a reasonable working assumption for tracking frontier progress. It is not a finding, and the sub-indices exist precisely because Epoch knows it.
What we would like to see is the residuals. For each model, which benchmarks does it beat its fitted curve on and which does it miss? A model that systematically outperforms on developer-reported tests and underperforms on third-party ones would be flagged by that table in a way no single score can flag it. The data and the code are public, so this is an afternoon's work for anyone who wants to try, and it would tell us more about the index than the index tells us about the models.
Sources
From the foundation