55 percent faster: reading the first Copilot RCT carefully
The GitHub and Microsoft controlled experiment found developers with Copilot finished a JavaScript HTTP server 55.8 percent faster. Reading notes on the design, the confidence interval nobody quotes, and what a single-task experiment can tell you about real work.
What was run
Sida Peng, Eirini Kalliamvakou, Peter Cihon, and Mert Demirer posted the write-up of GitHub's Copilot experiment on February 13. It is the first randomised controlled trial of an AI coding assistant we are aware of, and the headline number, a 55.8 percent reduction in task time, is going to be quoted in every sales deck this year. So it is worth reading what the experiment actually was.
The team sent 166 offers through Upwork and 95 developers accepted. They were randomised into a treatment group of 45 with access to Copilot, after a one minute instructional video, and a control group of 50 who could use anything else, including search and Stack Overflow. The task was to write an HTTP server in JavaScript as quickly as possible. Of the 95, 35 in each group finished both the task and the exit survey.
The number and its interval
The treatment group averaged 71.17 minutes. The control group averaged 160.89 minutes. That is the 55.8 percent, with a p-value of 0.0017. The effect is real in the sense that it is very unlikely to be noise. What gets dropped from the summaries is the 95 percent confidence interval, which runs from 21 percent to 89 percent. The experiment is consistent with Copilot saving a fifth of the time and with it saving almost nine tenths.
An interval that wide is what you get from 70 completers on a task with high variance in individual speed. It does not make the study wrong. It does mean that anyone who repeats 55 percent as if it were a measured constant is reporting the point estimate and hiding the uncertainty. The honest one-line summary is that Copilot made a scripted task substantially faster, by an amount the experiment cannot pin down to better than a factor of four.
Success rates are the other number worth knowing. The treatment group finished the task at a rate 7 percentage points higher than control, with a confidence interval from minus 11 to plus 25. That is a null result. Copilot made the people who completed the task faster, and had no measurable effect on whether they completed it.
Who benefited
The paper reports heterogeneous effects, which are the most interesting part and also the least reliable, because subgroup analyses on 70 people are underpowered by construction. The pattern they report is that less experienced developers gained more, that participants aged 25 to 44 gained more than younger or older ones, and that people who code more hours per day gained more. The authors read the first of these as promise for people entering software careers.
We would read it more cautiously. Less experienced developers on Upwork completing a standard task are the population most likely to be slowed by looking things up, which is exactly what Copilot replaces. The result says that autocomplete for boilerplate helps people who would otherwise search for the boilerplate. It does not say much about a senior engineer working in a codebase they know.
What a single-task RCT can and cannot say
The authors are explicit that this is a standardised programming task in an experiment rather than real collaborative work, and that the study does not look at code quality at all. Both caveats deserve more weight than they will get. An HTTP server in JavaScript is a task with hundreds of near-identical examples in the training data. It is the best possible case for a model that completes code from context. Nobody's job consists of writing that server.
The quality omission is the one that bothers us more. Time to completion measures how fast a working server appeared. It does not measure whether the code has a subtle bug, whether the developer understood what they accepted, or how long the next person spends reading it. A tool that produces plausible code faster could easily net out negative on a real project if review and debugging time rises, and this design cannot see that.
There is also the question of what the control group was doing. They had search and Stack Overflow, which is the right comparison for 2022. But the control condition still requires reading, judging, and adapting an answer, and the treatment condition often does not. Some of the speedup is the model being good. Some of it is the task being one where judgment was not required and skipping it was safe.
What we would want next
The obvious follow-up is the same design on a task inside an existing codebase, with a quality review by people blind to condition, and with enough participants to shrink that interval. The second follow-up is a field study that measures time to merged pull request rather than time to a passing script, since merge is where the review cost shows up.
Until those exist, the right way to cite this paper is with the interval attached and the task named. Copilot made an HTTP server appear somewhere between 21 and 89 percent faster in the hands of freelancers. That is a genuinely useful finding, and it is a much narrower one than the number that will circulate.
Sources
From the foundation