From Devin to 4 percent of GitHub commits: the agent flood and the slop backlash
Two years after Devin launched on a 13.86 percent SWE-bench score, one coding agent is estimated to author 4 percent of public GitHub commits, curl has shut its bug bounty and tldraw auto-closes outside pull requests. The bottleneck moved from writing code to reviewing it.
Where the curve started
On March 12, 2024, Cognition came out of stealth with Devin, described as an AI software engineer, and a headline number: 13.86 percent of SWE-bench issues resolved end to end without human help. The comparisons at the time were Claude 2 at 4.80 percent, GPT-4 at 1.74 percent and SWE-Llama-13b at 3.97 percent, and those baselines had been given hints about which file to edit. Devin had a shell, an editor and a browser in a sandbox, made a plan from a natural language request, and reported progress while it worked. It was non-public, the underlying models were undisclosed, and the company had raised $21 million.
We wrote at the time that the number was less interesting than the shape of the product. A model that could open a terminal and iterate was a different object from a model that completed lines. Two years on the number looks quaint and the shape looks like the whole industry. The question we want to take up is what happens when that shape works well enough that its output arrives faster than anyone can read it.
How much code is this
SemiAnalysis published an estimate on February 5 that 4 percent of public GitHub commits are currently authored by Claude Code, with a projection of more than 20 percent of daily commits by the end of 2026 if the trajectory holds. The figure comes from GitHub data run through their own model, and the piece does not lay out the method, the time window, or how Claude Code commits were distinguished from other tools. We are citing it because it is the number everybody is repeating this week, and we want to be explicit that it is one firm's estimate for one tool, not a measurement of AI-authored code overall.
Even discounted, the direction is not in doubt. And the 4 percent, whatever its precision, is concentrated. A commit to a personal repository nobody reads costs nothing. A pull request to a project with three volunteer maintainers costs one of them an hour. The projects that are shutting doors this month are the ones where the ratio of arriving contributions to available review hours went past the point where a person could keep up.
Three projects close a door
On January 15, Steve Ruiz posted a contributions policy for tldraw announcing that pull requests from external contributors would be closed automatically. The reasons listed were a surge of AI-generated pull requests with incomplete or misleading context, misunderstandings of the codebase, and minimal follow-up from authors. Issues, bug reports and discussions stay open, and a maintainer can reopen a closed PR that looks worth considering. Ruiz framed it as temporary, pending better tooling from GitHub, and noted it reverses a decision from 2021 to open the project to public code contributions.
On January 26, Daniel Stenberg announced that curl's bug bounty, which had run through HackerOne since April 2019, would end on January 31. BleepingComputer reported that the project received seven HackerOne reports in a 16-hour period in one week of January and 20 submissions by mid-month, following a steep rise through 2025. Stenberg wrote that "the never-ending slop submissions take a serious mental toll to manage and sometimes also a long time to debunk", and that the goal of ending the bounty was "to remove the incentive for people to submit crap and non-well researched reports". Security reports move to GitHub without payment from February 1. One example he gave was a report about a supposed HTTP/3 stream dependency cycle, complete with debugger sessions and register dumps, that referenced a function which does not exist in curl.
Then on February 3, The Register reported that GitHub itself had opened a community discussion, posted by product manager Camilla Moraes, on the rising volume of low-quality contributions. Options under consideration include letting maintainers disable pull requests entirely or restrict them to collaborators, deleting pull requests from the interface, finer-grained permissions for creating and reviewing PRs, AI-based triage, and some way of signalling that an AI tool was used. In the same discussion, Xavier Portilla Edo of Voiceflow estimated that one in ten AI-created PRs met the required standard. Jiaxiao Zhou of Microsoft described the review trust model as broken, because "reviewers can no longer assume authors understand or wrote the code they submit".
The review trust model was the load-bearing part
Zhou's sentence is the one we keep returning to, because it names the thing that actually broke. Code review in open source never worked by reading every line with fresh eyes. It worked because a submitted patch was evidence that a person had understood the problem, run the code, and was willing to answer questions about it. The reviewer checked the author's reasoning as much as the diff. When the diff arrives without any of that behind it, the reviewer has to rebuild the reasoning from scratch, and that is more work than writing the patch was.
A concrete example of the asymmetry. An agent produces a 400-line pull request that refactors error handling across a module, with a plausible description and passing tests. To merge it, a maintainer has to confirm the refactor preserves every behaviour, including the ones the tests do not cover, and that the description is accurate. That is a two-hour job for someone who knows the module. The agent spent four minutes. Multiply by the twenty such PRs that arrive in a week and the arithmetic explains tldraw's decision without needing anyone to be acting in bad faith.
Matthew Isabel at GitHub pushed back on counting AI PRs as the metric, on the grounds that a bad PR is bad regardless of where it came from. He is right about the individual PR and wrong about the system. A bad PR from a human costs that human something to produce, which caps the rate. A bad PR from an agent costs nothing, and the cap is gone. The problem is volume at zero marginal cost, and the identity of the author matters only because it predicts the volume.
What gating could look like
The remedies in GitHub's list split into two kinds. Some reduce volume: closing PRs, collaborator-only submission, deleting rather than closing. Those work, as tldraw shows, at the cost of the open contribution model that made these projects what they are. The others try to restore the evidence that review used to rely on. Attribution tells the reviewer what they are looking at. Finer permissions let a project require a track record before a PR is even created. Triage tries to reproduce, cheaply, the first pass a human used to do.
What we think is missing from the list is the thing curl was implicitly asking for. Stenberg's complaint is that submissions arrive without the work having been done. The fix is to make the work legible. A PR that ships with a reproduction, a description of what was run and what was observed, and a contributor willing to respond to questions within a day is reviewable whether or not an agent wrote the code. A PR that ships with none of that is not, whoever wrote it. Projects could state that as a requirement, enforce it mechanically for outside contributors, and let agents that meet the bar through.
The experiment we would want to see is a project running that policy for six months and publishing the numbers: PRs received, PRs meeting the bar, merge rate for each, and maintainer hours per merged PR. If the bar restores the review economics without closing the door, that is the answer. If it does not, then the pull request as an open institution may simply not survive contact with agents, and it would be better to know that from data than from watching one project after another close.
Sources
- VentureBeat: Cognition launches Devin (March 2024)
- SemiAnalysis: Claude Code is the inflection point
- LWN: Stenberg on the end of the curl bug-bounty program
- BleepingComputer: curl ending bug bounty after flood of AI slop reports
- tldraw: Contributions policy (issue #7695)
- The Register: GitHub ponders kill switch for pull requests to stop AI slop
From the foundation