The life cycle of SWE-bench Verified
OpenAI, which helped build SWE-bench Verified, has stopped reporting it after auditing the tasks its own model failed. The pattern, useful, then a target, then a contamination fingerprint, is the normal life of a coding benchmark and worth writing down.
What OpenAI reported
On 23 February OpenAI's frontier evals team said it will no longer evaluate models on SWE-bench Verified. The reasons, as reported, come in two parts. The first is test quality. OpenAI audited the problems that o3 failed across 64 runs, 138 problems or about 27.6 percent of the 500 task set, and found that 59.4 percent of those had serious flaws in their tests. Some tests demanded an exact function name that the task description never mentioned. Others checked behaviour pulled from the original pull request that was unrelated to the issue being fixed.
The second reason is contamination. The audit reported evidence that every frontier model examined, GPT-5.2, Claude Opus 4.5 and Gemini 3 among them, had been exposed to benchmark material in training. The mechanism is not mysterious. SWE-bench Verified is 500 public GitHub issues with public fixes, and GitHub is in every pretraining crawl. The answer sheet ships with the exam.
We have not been able to read OpenAI's own post, which returns an error from outside, so the figures above come from secondary reporting of it and we are treating them accordingly. The direction is not in doubt. The organisation that co-created the benchmark, in 2024, and used it in every model release since, has said the number no longer measures what it was built to measure.
Stage one: a useful instrument
SWE-bench Verified was a good idea when it appeared. The original SWE-bench pulled real issues from real repositories and asked whether an agent could produce a patch that passed the repository's tests. The problem was that a lot of those issues were underspecified or had tests that no reasonable patch could satisfy, so the score was capped by task quality rather than model ability. Verified was the human screened subset meant to fix that. The irony of this month is that the screen was not tight enough, and the failure mode it was built to remove is what the audit found.
For a while the benchmark did its job. It separated models. It rewarded agents that could read a codebase, run tests, and iterate. The scores climbed from the low thirties for GPT-4o in August 2024 to above 70 for leading agents by the middle of last year, and Claude Opus 4.5 is reported at 80.9 percent. That is the sort of curve that either means enormous progress or a benchmark that has been learned, and it is usually some of both.
Stage two: a target
Once a benchmark is in every model card it becomes a training objective, whether anyone intends it or not. Labs build agent scaffolds and tune them against the benchmark, because that is the number that will be compared. Test time budgets go up because more attempts pass more tests. Prompts get refined against the failure cases. All of that is legitimate engineering, and all of it moves the score without necessarily moving the ability the score was supposed to stand for.
SWE-bench Pro, the replacement OpenAI now points to, is built to resist this for a while. Its 1,865 tasks average 107 changed lines across 4.1 files, with a 10 line minimum and more than 100 tasks needing over 100 lines, against 161 tasks in Verified that need a one or two line change. Top models score in the twenties to high fifties on Pro, with Kimi K2.6 reported at the top at 58.6, against 70 to 80 on Verified. Same models, different instrument, a 20 to 50 point gap. That gap is a measurement of how much of the Verified score was the instrument.
Stage three: a contamination fingerprint
The third stage was visible last May, which is what makes this month's news a confirmation rather than a surprise. The SWE-rebench paper from Nebius built an automated pipeline that pulls fresh tasks from GitHub continuously, over 21,000 Python tasks, so that models can be evaluated on issues created after their training cutoff. They then compared scores on the fresh tasks against scores on SWE-bench Verified.
The comparison is the fingerprint. DeepSeek-V3-0324 scored 39.7 on Verified and 21.3 on fresh SWE-rebench tasks from March and April 2025. The earlier DeepSeek-V3-1226 scored 35.2 on Verified and 21.9 on the fresh tasks. Two checkpoints that perform the same on unseen problems but diverge on the old benchmark, with the newer one ahead, is exactly what you would see if the newer checkpoint had seen more of the old benchmark. Qwen2.5-72B-Instruct, by contrast, was at 11.3 and 9.3, roughly consistent. The authors put it carefully, saying the divergence may suggest contamination, and we think that care is right, but the pattern is what contamination looks like.
At this stage a benchmark has one remaining use. It tells you how much a model has memorised, by comparison with a fresh set. That is a real measurement, and it is the only one SWE-bench Verified can still make.
What to do with the next one
Every static coding benchmark will go through these three stages, and Pro will too. The only design that avoids stage three is the one SWE-rebench uses, continuous collection of tasks dated after the model's cutoff, with contamination flagged by date rather than argued about. That costs ongoing effort and produces a moving target, which makes year over year comparison harder. We think that trade is worth it, because the alternative is a fixed target that stops meaning anything after about eighteen months.
The practical request we would make of anyone reporting a coding score this year is to report two numbers, one on the fixed benchmark and one on a time split set of tasks created after the model's training cutoff. The gap between them is the most informative single figure available, and SWE-bench Verified, in its last useful act, has shown roughly how large it can get.
Sources
From the foundation