Regulators and courts price open weights: the GPAI Code and Bartz v. Anthropic
The EU's Code of Practice exempts most open-source models from its documentation rules, and Anthropic's $1.5 billion settlement pays for how books were downloaded rather than for training on them. Together they sketch a safe path for anyone releasing a model.
What the EU decided
The Commission published the General-Purpose AI Code of Practice on July 10, and the AI Act's obligations for general-purpose model providers took effect on August 2. The Code has three chapters. Transparency and Copyright apply to every provider of a general-purpose model under Article 53. Safety and Security applies only to providers of models with systemic risk under Article 55, which the Act defines by a training compute threshold of 10^25 FLOPs. Twenty-one organisations have signed so far, including Amazon, Anthropic, Google, Microsoft, OpenAI and Mistral, with xAI signing only the safety chapter.
The open-source carve-out lives in Article 53(2). Providers who release a model under a free and open-source licence, one that lets users access, use, modify and redistribute it with attribution, are exempt from the technical documentation requirements in 53(1)(a) and (b). They are not exempt from the copyright policy obligation or from publishing a sufficiently detailed summary of training data. And the exemption disappears entirely for a model above the systemic-risk threshold, where every obligation applies regardless of licence.
So for a research lab the regulatory position is now legible. Release under a genuinely open licence, stay below 10^25 FLOPs, publish a training data summary and a copyright policy, and the rest of the compliance burden falls away. A restricted licence of the Llama kind does not qualify, which is one more place where the definition of open now has a price attached.
What the court decided
Bartz v. Anthropic is the first case to rule on fair use for language model training, and the ruling split. On June 23 Judge William Alsup in the Northern District of California held that training on books was fair use, calling it exceedingly transformative. He also held that Anthropic buying print books and scanning them destructively into a digital library was fair use. What he did not accept was Anthropic's download of more than seven million books from Books3, Library Genesis and Pirate Library Mirror to build a permanent internal library. That, he ruled, was not protected, and he set a trial on the pirated copies and the resulting damages, including for willfulness.
Anthropic settled before that trial. The settlement is $1.5 billion, covering around 500,000 titles out of the seven million copies downloaded, which works out to roughly $3,000 per work before fees. The Authors Guild page explains eligibility: works registered with the Copyright Office within three months of publication or before August 10, 2022, with an ISBN or ASIN, and a claims deadline of March 30, 2026. The Guild, for its part, said in June that it expected the fair-use finding to be reversed on appeal. With a settlement, that appeal will not happen in this case.
The number to keep in mind is that the $1.5 billion was paid for acquisition, not training. Under the June ruling, if Anthropic had bought and scanned every one of those books, there would have been no liability at all. The liability came from where the files were downloaded from.
What the two decisions say together
Read side by side, the regulator and the court are pointing at the same set of practices. The EU wants a training data summary and a copyright policy. Alsup wants lawfully acquired inputs. Neither is asking for anything a careful lab should not already be doing, and both are punishing the same shortcut, which is treating a shadow library as a data source.
For open-model releasers specifically, that is good news with a condition. The open licence buys regulatory relief in Europe, and the training-is-fair-use finding, while it is a single district court and could be reversed elsewhere, means that publishing weights trained on lawfully obtained text is not itself the risk. The risk is a provenance chain that runs through LibGen, and an open release makes that chain easier for a plaintiff to find, not harder.
There is a real asymmetry here that we do not want to gloss. A closed lab that used pirated books has to be caught. An open lab that publishes its data manifest has already confessed. The response to that asymmetry cannot be to hide the manifest, because the EU now requires a summary anyway and because a hidden manifest defeats the purpose of releasing the model. The response has to be a manifest with nothing in it that needs hiding.
The safe path, as far as we can see it
Here is what we would tell a group planning an open release this autumn. Use a licence that actually meets the free and open-source test, with no field-of-use restrictions, or accept that you get none of the Article 53(2) relief. Keep a per-source manifest with licence and acquisition method for every component, and be able to show that no component came from a piracy mirror. Publish the training data summary the Act requires, in a form that a rights holder could search. Stay below the systemic-risk threshold unless you are prepared for the full safety chapter.
None of this is legal advice and all of it rests on one court and one Code. Alsup's ruling binds nobody outside his courtroom, and other judges have the Kadrey v. Meta record in front of them with Meta's own decision to train on pirated material. The Code is voluntary and its interpretation will shift as the AI Office starts enforcing. But the direction is consistent, and it is the direction open science was already going.
What we would like to see next is a study of the actual cost of a clean corpus. Anthropic spent $1.5 billion because it did not buy the books. Someone should estimate what buying and scanning them would have cost instead, at the scale a frontier model needs, and whether the quality difference between a bought library and a pirated one is even measurable. Our guess is that the clean path is cheaper than the settlement and slightly worse on benchmarks. We would like to be corrected with data.
Sources
From the foundation