The Stack v2 and the opt-out: building a code dataset people can leave
BigCode built a code corpus roughly four times the size of the one behind StarCoder, sourced from Software Heritage, and ran an opt-out process before training. What that process looked like and what it removed.
What was built
The BigCode project published the StarCoder 2 paper today, together with The Stack v2, the dataset the models were trained on. The dataset is about four times the size of the one used for the original StarCoder and it was assembled in partnership with Software Heritage, the non-profit archive of public source code. The models come in 3B, 7B and 15B sizes, released under the OpenRAIL licence, with 66 authors on the paper.
The numbers give a sense of the scale. The raw pull came from the 6 September 2023 snapshot of the Software Heritage graph, covering 619 programming languages. After deduplication the source code portion for the 32 major languages they report on came to around 6,457 GB across 784 million files. The training mixes came to 622 billion unique tokens for the 3B model, 659 billion for the 7B, and 913 billion for the 15B, with total training running between 3.3 and 4.3 trillion tokens over no more than five epochs. Alongside the code, the team pulled in GitHub pull requests, Kaggle notebooks and code documentation.
What interests us about this release is less the benchmark table and more the pipeline. This is the largest openly documented attempt we know of to build a training corpus where the people whose work is inside it were given a way to check and a way to leave.
Why Software Heritage changes the sourcing
Most code datasets are built by crawling GitHub directly. The Stack v2 instead sits on top of the Software Heritage archive, and the paper credits the archive as a partner rather than just a source. The practical consequence is provenance. Every file in the dataset can be traced back to a Software Heritage persistent identifier, a SWHID, which points at a specific content hash inside the archive. If you want to know where a training example came from, there is a stable answer.
The extraction kept only the latest commit on each repository main branch, then deduplicated on content hash. Licence handling sorted files into three buckets, permissive, non-permissive copyleft, and unlicensed. The team kept permissive and unlicensed files and excluded copyleft and commercial licences. That is a policy choice that will not satisfy everyone, and the inclusion of unlicensed files in particular is the kind of decision we would expect to see argued about. But it is written down, which is more than most corpora offer.
How the opt-out ran
The governance mechanism is the tool called Am we in The Stack. A developer can look up whether their repositories are in the dataset and file a request to have them removed. For v2 the team announced the opt-out window on X with a deadline of 20 November 2023, then processed the requests before training began.
The result of that process is stated precisely in the paper. 1,561 repositories associated with 91 users and organisations were eliminated, removing 22,066 files from the source code dataset. Set against 784 million files, that is a rounding error in size. Set against the alternative, which is a corpus nobody can leave, it is the whole point.
Personal information was handled separately. The team ran the StarPII model over the code to redact names, emails, keys, passwords, IP addresses and usernames. In the pull request and discussion data, usernames were replaced with per-conversation participant counters so that a thread stays readable without naming the people in it.
What the small number tells you
Ninety-one opt-outs from an archive covering a large fraction of public code is a small response and we think it deserves an honest reading rather than a flattering one. It could mean that most developers are content to have their permissively licensed code used for training. It could equally mean that most of them never saw the announcement, which went out on a single social network with a deadline a little over two months before the paper. An opt-out process only measures consent among the people who found out about it.
The mechanism also runs after the fact. A file is in the archive by default and leaves only when someone asks. An opt-in design would produce a much smaller corpus and probably a worse model, and we do not think BigCode was wrong to choose the design they did. But the choice should be named as a choice. The consent here is of the kind where silence counts as yes.
The part that does hold up regardless of those criticisms is traceability. Because every file carries a SWHID, a future opt-out can be honoured precisely, and an outside auditor can check what was and was not included. That is a property most datasets in this field do not have and it does not depend on how many people used the form.
What we would want next
The obvious follow-up experiment is a retraining study. Take the 3B model, remove a further random sample of repositories equal in size to the opt-out set, retrain, and report whether the benchmark scores move. If they do not, and we suspect they will not at this scale, then the cost of honouring opt-outs is close to zero and there is no excuse for other groups to skip the process.
The harder question is what happens when the opt-out set stops being tiny. A corpus that can be left is only meaningful if leaving is easy and widely known, and if a lot of people did leave, the licence filtering and the deduplication would both need to be rerun. The Stack v2 gives us the first pipeline where that rerun is even possible. We would like to see it exercised before the next version rather than after.
Sources
From the foundation