SWE-Lancer prices the benchmark in dollars
OpenAI’s SWE-Lancer takes 1,488 real Upwork tasks from the Expensify repository, worth $1 million in actual payouts, and scores a model by how much it would have earned. Claude 3.5 Sonnet earns about $403,000. What a dollar metric captures that pass rates hide, and where it misleads.
The benchmark
OpenAI released SWE-Lancer on Monday. The idea is to stop asking what fraction of coding tasks a model can solve and ask instead how much money it would have made solving them. The tasks are 1,488 real freelance jobs posted by Expensify on Upwork, drawn from Expensify's open-source repository, with the actual prices Expensify paid, from $250 up to $32,000 per task and $1 million in total. A public subset called SWE-Lancer Diamond has 502 tasks worth $500,800.
There are two task types. The 764 individual contributor tasks ask the model to fix the bug or build the feature, and the result is checked by end-to-end tests that drive the application in a browser with Playwright. The 724 manager tasks ask the model to choose the best of several proposed solutions, which is graded against the proposal that was actually chosen. OpenAI contracted 100 professional engineers to write the tests, and each test went through three rounds of validation for quality, coverage and fairness.
The numbers
On the full set, Claude 3.5 Sonnet passes 21.1 percent of individual contributor tasks and 47.0 percent of manager tasks, for a combined 33.7 percent and earnings of about $403,000. o1 at high reasoning effort passes 20.3 and 46.3 percent for $380,000. GPT-4o passes 8.6 and 38.7 percent for $304,000. On the Diamond subset Sonnet earns $208,000, with 26.2 percent on implementation tasks and 44.9 percent on manager tasks.
Two patterns are consistent across models. Manager tasks are much easier than implementation tasks, which fits the general finding that judging a solution is easier than producing one. And more attempts help a great deal: for o1, six additional attempts at each task nearly triples the number solved, and raising reasoning effort from low to high moves pass@1 on implementation tasks from 9.3 to 16.5 percent.
The failure analysis is the most useful part of the paper for anyone building agents. The authors observe that models pinpoint the source of an issue remarkably quickly and then fix the wrong thing, because they have a limited understanding of how the issue spans multiple components. Localisation is close to solved. Root cause is not.
What a dollar metric gets you
A pass rate treats every task as equal. A dollar metric weights tasks by what someone was willing to pay, which is a market's estimate of how much skill and time the task takes. The authors argue that harder tasks pay more, especially those that need specialised knowledge or sat unresolved for a long time, so the payout gives you a difficulty gradient without anyone hand-labelling difficulty.
This does two things a pass rate cannot. It makes the number legible outside the field, because a product manager knows what $400,000 of freelance work means and does not know what 33.7 percent of a benchmark means. And it makes the gap between models economically interpretable: the difference between Sonnet and GPT-4o on this set is about $100,000 of work, which is a more useful sentence than ten percentage points.
It also exposes something a pass rate hides. If a model's earnings are concentrated in the cheap tasks, then a respectable pass rate can coexist with a small dollar total, and vice versa. Reporting both numbers, as the paper does, lets a reader see whether a model is solving many small things or a few big ones.
Where the dollar misleads
Price is not difficulty. A task's payout on Upwork reflects urgency, how long it sat open, how specialised the required knowledge is, and how Expensify sets its bounties, on top of how hard the work is. A $32,000 task is not 128 times harder than a $250 one. So a model that clears one big bounty by luck or by having seen a similar issue in training can post a large earnings jump that a pass rate would show as a single task.
The single-repository design is the other limit. Every task comes from one company's codebase, with its particular architecture and test culture. The authors are open about this, and about the fact that freelance tasks are more self-contained than the average full-time engineering ticket, that infrastructure and DevOps work is under-represented, that the evaluation is text-only and drops the screenshots and videos a human freelancer would get, and that some issues were public on GitHub during 2023 and 2024 and may have been in training data. A dollar figure lends the result a precision that the sampling frame does not support.
There is also a subtler problem with reading earnings as value. The tests check whether the fix passes end-to-end, which is what Expensify would have paid for. They do not check whether the change is one a maintainer would merge. An agent that earns $403,000 by passing tests with code a reviewer would reject has earned it under the benchmark's rules and would not have been paid in practice.
What we would do with it
We like the metric and we would report it alongside, never instead of, a pass rate and a per-price-band breakdown. The band breakdown is the piece that resolves most of the objections above: if a model's pass rate is flat across price bands, then dollars and percentages agree and the dollar total is a fair summary. If the pass rate collapses above $1,000, the dollar total is being carried by volume and the reader should know.
The broader move, benchmarks built from tasks that someone actually paid for, is one we would like to see repeated in other domains. It is expensive, it requires a cooperating company, and it produces a test that can say something about economic value rather than about a synthetic proxy. The next version should draw from several repositories, and we would want to see the manager task graded against more than one human's choice.
Sources
From the foundation