What LAION-5B is, and is not

LAION-5B is not a collection of images. It is a table of about 5.85 billion rows, each with a URL, the alt text that was next to the image in the Common Crawl snapshot, a hash, a language tag and a score from a safety classifier. The images themselves stay wherever they were on the web. Anyone who wants to train on the dataset downloads them from the original hosts. That structure was chosen to limit size, liability and copyright exposure, and it matters for everything that follows, because it means the dataset can be audited without anybody having to distribute the images.

The English subset, LAION-2B-en, was the main training data for Stable Diffusion 1.5. The 2.0 release filtered out anything the classifier scored as unsafe above 0.1, which made the model much worse at explicit content and left 1.5 as the community's preferred base for that purpose. The 2.1 release trained 2.0 further on both safe and moderately unsafe material. Google's Imagen used the earlier LAION-400M alongside internal data, and Google's own audit of that dataset found enough inappropriate content that it was judged unfit for public use.

How the Stanford team searched

David Thiel's report describes work that started in September 2023 and went through several approaches. Keyword search on translated captions was tried first and abandoned, partly because machine translation of non-English captions produced garbage, and partly because the initial PhotoDNA hits mostly had generic captions that gave no hint of what the image was. Text turned out to be nearly useless for finding this material.

The method that worked was to take every entry the LAION safety classifier scored above 0.995, about 32 million items, and send the URLs to Microsoft's PhotoDNA, which matches perceptual hashes of known abuse imagery. Matches went to the Canadian Centre for Child Protection for human verification through its Arachnid API. Verified matches were then used as seeds for nearest neighbour queries over the CLIP embeddings that LAION distributes, and those neighbours were checked again with PhotoDNA and with a classifier from Thorn. In parallel the team joined the datasets' MD5 hashes against hash lists from NCMEC, which catches images that have gone offline but only if they are byte-identical.

No abuse imagery was downloaded or stored. Candidates from the neighbour search that had no hash match were pulled to a RAM disk on a temporary VM, run through the classifier, and deleted, with only URLs of matches passed to the verifiers.

What they found

The headline figure is 3,226 dataset entries of suspected CSAM. Breaking that down, the classifier cutoff plus PhotoDNA gave 1,679 hits, of which 689 were validated by C3P as CSAM or likely CSAM. Nearest neighbour searches added 162 PhotoDNA hits and the MD5 join added 229 unique matches that PhotoDNA had not seen. The Thorn classifier on the neighbour sets surfaced 989 candidates with 183 validated. About 30 percent of the URLs sent to PhotoDNA were already dead, and the report is explicit that the totals are a significant undercount for that reason and because hash sets are incomplete.

Some images appeared up to eight times in the dataset. The report flags that repetition as its own harm, since repeated training examples are the ones diffusion models are most likely to reproduce closely. Hits came from the content delivery networks of Reddit, Twitter, Blogspot and WordPress as well as from adult sites, which is a reminder that the dataset simply reflects what the web hosted in 2021 and 2022.

LAION pulls the datasets

On 19 December, the day before the report's initial release, LAION published a note saying it was temporarily taking down its datasets to ensure they were safe before republishing them. The note pointed to filtering measures described in 2021, said the organisation was working with the Internet Watch Foundation, invited the Stanford researchers to help improve its filters, and cited its obligations under GDPR's right to erasure.

That last point deserves a closer look. The report notes that LAION's crawl did try to discard images that were both NSFW and matched an underage filter, but the keyword list behind that filter was thin and included typos, and the images were never checked against known CSAM hash lists at all. Partnering with NCMEC, Microsoft, C3P or Thorn at compilation time would have caught most of what Stanford later found. LAION clearly intended to filter this material. The gap was that nobody with the right hash lists was in the room when the dataset was built.

The part that is uncomfortable for open science

Every step of Stanford's method depended on the dataset being public. The URLs, the safety scores, the hashes and the embeddings that seeded the neighbour searches were all distributed by LAION as part of the release. A closed training set at a commercial lab could contain exactly the same material and nobody outside the lab would ever be able to check. The report's recommendations to dataset builders assume they will run these checks themselves, and for closed datasets we have to take that on faith.

So the finding cuts in two directions. Open release put a dataset with abuse material in it into thousands of researchers' hands, and the report says that removing it from those downloaded copies is one of the hardest steps because nobody knows who has a copy. Open release is also the only reason the problem was found, counted, reported to NCMEC, and fixed at the source. We do not think the answer is to stop releasing datasets. We think the answer is that hash checking against known CSAM lists has to be a precondition of release, the way a licence file is, and that the organisations holding those hash lists need to make them available to non-profit dataset builders.

What happens to the models

Removing rows from a table is easy. Removing what a model learned from those rows is not. The report is candid that even if the embeddings for known matches were removed from a model, nobody knows whether that would change its ability to produce the material. The more general tools are concept ablation and concept erasure, which retrain a model to discard a concept, but the report points out that erasing this particular concept may require access to examples of it, which is not a training run anyone should be doing.

What we would want to see in the next year is a republished LAION with a documented hash-screening pass, a public statement from Stability AI about what Stable Diffusion 1.5 was trained on, and hosting platforms adopting a reporting path for models known to produce this material. The auditing method is now written down. The question is whether anyone building the next dataset uses it before release rather than after.

Sources

  1. David Thiel, Identifying and Eliminating CSAM in Generative ML Training Data and Models, Stanford Internet Observatory, December 2023
  2. LAION, Safety review for LAION 5B, 19 December 2023