The number

Carlos Jimenez, John Yang and colleagues at Princeton posted SWE-bench on October 10. It contains 2,294 tasks drawn from real GitHub issues in twelve Python repositories, and the best result in the paper is Claude 2 resolving 1.96 percent of them in the realistic setting. With an oracle that hands the model exactly the files the human fix touched, Claude 2 reaches 4.8 percent and GPT-4 1.7 percent. ChatGPT-3.5 manages 0.2 percent. Those are the numbers at launch, and we want to record why they are the right kind of low.

Coding benchmarks before this were function-sized. HumanEval asks for a few lines against a docstring. Models had long since made that easy, and a saturated benchmark tells you nothing about the difference between two models that both pass it. SWE-bench was designed to be hard enough that nobody passes and cheap enough that anyone can check.

How the tasks were chosen

The construction is a filter with three stages, and the stages are the reason the benchmark works. The authors scraped about 90,000 pull requests from twelve popular repositories, including Django, scikit-learn, Matplotlib, Sphinx and SymPy. They kept only the merged pull requests that resolve a linked issue and also change the repository's test files, on the theory that a contributor who added a test was checking their own fix. Then they ran the tests before and after each candidate patch and kept only the ones where at least one test flipped from failing to passing, discarding anything that failed to install or run.

That last stage is the whole design. Every task in the final set comes with a fail-to-pass test that the reference fix makes pass, plus a set of tests that must keep passing. The median task has 51 of those. Grading a model's patch means applying it and running the tests, with no human in the loop and no judgement call. Django alone contributes 850 tasks and SymPy 386, so the distribution is lumpy, which the paper is open about.

What the tasks look like

The averages describe a maintainer's job rather than a puzzle. The issue text averages 195 words. The codebase averages 3,010 files and 438,000 lines. The reference patch edits 1.7 files, 3 functions and 32.8 lines, and at the tail a single gold patch changes 5,888 lines across 31 files. A model has to find the right place in a large repository from a short description, make a change that fits the surrounding code, and not break anything else. Nothing in the input says where to look.

Because the codebases dwarf any context window of the time, the authors used BM25 to retrieve files for the model, at limits of 13,000, 27,000 and 50,000 tokens. More context made things worse. Claude 2 went from 1.96 percent at 13,000 tokens to 1.22 percent at 50,000. The retriever also often missed the right files entirely, excluding all of the oracle files in over half of instances at the middle limit. Some of the gap between 1.96 and 4.8 is retrieval and some is the model, and the paper is careful to separate the two.

How the models failed

The failure analysis is the part of the paper we would hand to someone building an agent today. Model patches that applied cleanly were short, 30.1 lines on average against 74.5 for the corresponding gold patches, and they almost always touched a single file. Models wrote primitive Python rather than using the libraries already in the repository. They fixed the literal symptom in the issue and left the structural problems a human maintainer would have addressed in the same pull request. Asking models to regenerate whole files instead of producing a patch made things worse, 2.2 percent against 4.8 for Claude 2.

The fine-tuned SWE-Llama models, 7B and 13B CodeLlama variants trained on 19,000 issue-pull request pairs from 37 other repositories, are a useful control. In the oracle setting they resolve 3.0 and 4.0 percent, close to Claude 2, and the 13B model and Claude 2 solve substantially different subsets of the tasks. With BM25 retrieval they fall to 0.7 percent, because they were trained on oracle context and could not cope with irrelevant files. Distribution shift between training context and evaluation context is a failure mode that every coding agent since has had to design around.

What the launch numbers set up

Three choices in this paper became defaults. Real repositories instead of synthetic tasks. Execution-based grading instead of similarity to a reference. A construction pipeline that can be re-run on new issues so the benchmark can be refreshed as models absorb the old ones. The 1.96 percent figure did something else. It gave the field a number low enough that any real progress would be visible, and the paper's own discussion says the next steps are agents that find their own context and tool use during generation, which is the direction the field took.

The thing we would want to check, if we were starting from this paper, is contamination. The repositories are public and heavily represented in pretraining data. The authors split tasks by whether the pull request predates 2023 and found little difference for most models, which is reassuring, but the benchmark is a fixed set and the models will change. A benchmark this good gets optimised against, and the pipeline that built it is the tool for building the next one.

Sources

  1. Jimenez et al., SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (arXiv 2310.06770)