What was audited

The Data Provenance Initiative posted its audit on October 25. Shayne Longpre, Robert Mahari, Anthony Chen, and fourteen co-authors from a mix of universities, companies, and Creative Commons traced more than 1,800 text datasets used for instruction and alignment fine-tuning back to their sources, their creators, and their licenses. The datasets are the ones that appear in the popular open fine-tuning collections, the ones everyone downloads to train a chat model, so the sample is not random and that is the point.

For each dataset they recorded where the text came from, who assembled it, what license the original creators attached, and what license the hosting page claims. The last two are the comparison that produces the headline numbers.

The two rates

More than 70 percent of the datasets, as listed on widely used hosting sites, had no license information at all. That is the omission rate. Among the datasets that did carry a license on the hosting page, more than 50 percent had one that did not match what the audit found at the source. That is the error rate. Together they mean that if you pick a fine-tuning dataset off a hosting site and read its license field, you are most likely to find nothing, and if you find something there is an even chance it is wrong.

The direction of the errors matters. The paper's phrasing is that licenses are frequently miscategorised, and the practical case is a dataset that inherits a non-commercial or share-alike condition from its source, gets reposted with a blank or with a permissive tag, and is then used as if it were free. Nobody in that chain necessarily acted in bad faith. Reposting is easy, license tracing is tedious, and the field's norms have rewarded the first and never asked for the second.

What the closed datasets look like

The audit also splits the collection by how restrictive the licenses are, and finds that datasets under closed or non-commercial terms are not spread evenly across task types. They dominate lower-resource languages, creative tasks, and synthetic data generated by models. The languages finding is the one we would flag for anyone building outside English. The datasets that exist for those languages are the ones most likely to carry conditions, so the commercial open-model ecosystem is effectively English-first by license as well as by volume.

The synthetic data point is a different problem. Data generated by a commercial model carries whatever restrictions the model's provider put in its terms of service, and those terms are not licenses in the Creative Commons sense. They sit outside the whole framework the audit uses, and the audit is honest that copyright and fair use interpretation differ by jurisdiction and that the paper is not legal advice for any of them.

What the Explorer fixes

The team released the Data Provenance Explorer at dataprovenance.org, an interface where you can filter the audited collections by source, by license category, by language, and by task, and download the subset that meets your conditions. If you need a commercially usable instruction set, you can now build one from audited components instead of hoping the hosting page is right.

That is a real improvement and it is worth being precise about what it does. It replaces one team's guess about each dataset's license with another team's careful reading of that dataset's license. The reading is better and the sample is large, and it is still a reading, of a document that may itself be ambiguous, applied to a dataset whose text may have come from somewhere the document does not mention. The audit gives you a defensible position. It does not give you certainty, and it does not cover the 1,800th dataset that was posted the week after the audit closed.

What it cannot fix

The deeper problem is that provenance decays. A dataset with a clean license gets merged into a collection, the collection gets filtered and deduplicated and reposted under a new name, and the license field on the new page is whatever the reposter typed. Every hop loses information and nothing in the current tooling carries it forward. An audit is a snapshot. It measures the loss at one moment and the loss resumes the next day.

The fix for that is infrastructure that makes the license travel with the data, and no number of audits substitutes for it. Hosting sites could require a source field and a license field before upload and refuse blanks. Collection builders could carry per-example provenance instead of per-collection. Neither is hard technically. Both require the field to decide that a dataset without a license is not a dataset, and the audit's contribution is to put a number on how far we are from that. Seventy percent is the distance.

Sources

  1. Longpre, Mahari, Chen et al., The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing and Attribution in AI (arXiv 2310.16787)