Getty v. Stability AI: the first big training-data lawsuit, and why it was about images
Getty Images announced this week that it is suing Stability AI in the High Court in London over the images used to train Stable Diffusion. The first major rightsholder to go to court over training data picked a diffusion model, and the reasons are about evidence as much as law.
What Getty said
On January 17 Getty Images published a short statement saying it had commenced legal proceedings against Stability AI in the High Court of Justice in London. The claim is that Stability AI unlawfully copied and processed millions of images protected by copyright, together with the associated metadata that Getty owns or represents, without a licence, for its own commercial benefit and to the detriment of the content creators. The statement is brief and avoids the word scraping, though that is what it describes.
The part of the statement that matters most for what comes next is about licensing. Getty says it has provided licences to leading technology companies for purposes related to training AI systems, and that Stability AI did not seek one. In Getty's phrasing, Stability chose to ignore viable licensing options. The statement also says Getty believes AI has the potential to stimulate creative work, and that its objection is to development that does not respect intellectual property rights. That framing is deliberate. Getty accepts that training on images can be done lawfully, and its complaint is that there was a price and Stability declined to pay it.
Why a diffusion model and not a text model
Large language models have been trained on scraped web text for years, and no major rightsholder has sued over it. The first big case is about a text to image model. We think there are three reasons, and none of them is that images are legally special.
The first is that the training set is public. Stable Diffusion was trained on LAION-5B, a dataset of more than five billion image and caption pairs released in March 2022 by the LAION nonprofit, built by extracting pairs from Common Crawl. Anyone can download the index and search it for their own images. A plaintiff does not have to guess whether its work was used or argue from statistical inference. It can look. Text model training sets are typically not published at that level of detail, which makes the first step of a copyright claim, showing that copying happened, much harder.
The second is that images carry their own evidence. A photograph from a stock agency is a discrete work with a registration, an owner and usually a visible watermark, and Getty's statement stresses that the metadata attached to each image was copied along with it. Metadata and watermarks are the kind of thing that survives a pipeline in recognisable form. There is no equivalent for prose. A language model that has absorbed a news article does not carry the newspaper's masthead around with it.
The third is that the output competes directly. Getty sells images. A model that generates images in response to a prompt is a substitute for that product in a way that a chatbot is not yet an obvious substitute for any particular publisher. When a plaintiff is going to argue market harm, it helps to be able to point at the market.
Scraping versus licensing
The argument Getty is setting up is about the existence of a licensing market. If a market for training licences exists, and Getty says it has already sold into it, then unlicensed training bypasses a transaction that was available. This is the version of the argument that fits a fair dealing or fair use analysis best, because it goes to market effect and to whether the copying was necessary.
The defence every lab building on web data relies on is that training learns statistical patterns from publicly accessible material and does not reproduce any particular work. Stability has not yet answered the claim in court, so we will not put words in its mouth. That is still the argument the case will have to test. The difficulty for it is that the LAION index makes the specific works identifiable, and a defence that works at the level of statistical abstraction has to survive a plaintiff who can name the files.
What is worth watching is the venue. Getty filed in London rather than in the United States, and the two systems treat copying for analysis differently. Which forum a rightsholder picks first tells you where it thinks the law is most on its side, and that choice will be studied closely by everyone deciding where to file next.
What this means for open datasets
The Getty statement does not mention LAION. But it is the reason the case exists in the form it does. An open dataset with a public index is what let Getty build its claim, and the same openness is what let researchers audit LAION for duplicates, for personal data and for harmful content. The property that makes a dataset verifiable is the property that makes its users suable.
That is the tension we expect to define the next few years. The open science argument is that training data should be published so results can be checked and problems can be found. This case is the first demonstration of the cost of that argument to the people who act on it, and the implicit lesson for a lab watching from the outside is to train on data nobody can inspect. We would rather that lesson not be the one that sticks, and the way to avoid it is to make the licensing question answerable in the open, with datasets whose provenance is documented well enough that the argument is about the licence and not about whether copying can be proved.
The question we want answered by the time this reaches a judgment is simple. Did the model copy the images, or did it learn from them, and does the law in the relevant jurisdiction think those are different things. Everything else in the AI and copyright debate follows from the answer.
Sources
From the foundation