A different kind of contamination

Most contamination discussion is about test sets. If the exact test examples are in the pretraining data, the score is meaningless, and people grep for them. Changmao Li and Jeffrey Flanigan, in a paper posted at the end of December and accepted at AAAI 2024, look at something broader they call task contamination. The training examples for a task, along with their labels, end up in the pretraining corpus, and the model has effectively been trained on the task before you show it your handful of demonstrations. The evaluation is then no longer zero-shot or few-shot in any meaningful sense, even if the specific test items are clean.

That is harder to detect than test set leakage because there is nothing exact to search for, and for closed models there is no corpus to search. The paper uses four methods. Training data inspection and task example extraction look for direct evidence. A membership inference attack checks whether the model reproduces exact outputs for generation tasks. The fourth method is the one we think everyone should copy, because it needs no access to training data at all.

The before-and-after-cutoff test

The idea is chronological. Every model has a date after which its training data was collected. Every dataset has a release date. If you evaluate a model on datasets released before its cutoff and on datasets released after, and the model does systematically better on the older ones, the simplest explanation is that it saw the older tasks during training. The authors ran this across twelve models, from the GPT-3 series (davinci through GPT-3.5-turbo) to open models including Fairseq MoE, GPT-J, OPT, BLOOM, LLaMA, Alpaca and Vicuna, on sixteen classification tasks and one semantic parsing task.

The datasets split into two groups. The pre-2021 group includes familiar names like RTE, SST-2, MRPC, QNLI, BoolQ and WiC. The post-2021 group includes StrategyQA, two NewsMTSC variants, NLI4Wills, CREPE, FOMC and NewsMet, released between 2021 and 2023. The GPT-3 series has training data cutoffs from October 2019 through September 2021, so the second group is after the cutoff for all of them.

The measure they use is whether the model beats the majority baseline, meaning a classifier that always predicts the most common label. That choice sidesteps the problem of averaging accuracy across datasets of different difficulty. Datasets released before the cutoff had a significantly higher chance of being beaten, in both zero-shot and few-shot settings, and the difference between the pre and post populations was significant at the 99 percent level by a Mann-Whitney U test.

The result that stings

The finding we keep coming back to is on the clean side of the cutoff. For classification tasks where task contamination was not possible, models rarely showed statistically significant improvement over the majority baseline, zero-shot or few-shot. Read plainly, when the model could not have seen the task before, a few demonstrations in the prompt mostly did not lift it above guessing the common label. The few-shot ability the field has been describing since 2020 looks, on this evidence, to be substantially a memory of tasks that were in the corpus.

There is a confound the authors acknowledge. Newer datasets may simply be harder, since people now build benchmarks that existing models fail. That would also produce lower post-cutoff performance without contamination. The paper's other methods are there to close that gap. On the Spider semantic parsing task they ran a membership inference attack across all models and found a correlation of 0.88 between the number of training examples extractable from a model and its accuracy on the task. For the GPT-3 series the number of extractable examples grew with each version from davinci to GPT-3.5-turbo, and tracked the rise in zero-shot performance on those tasks.

Why the chronological test is worth adopting

The method is cheap. You need release dates for your datasets, cutoff dates for your models, and an evaluation you were going to run anyway. It works on closed models where you cannot inspect anything. The authors are clear about its trade-off. Direct methods like example extraction have high precision and low recall, because they only catch what you happen to search for. The chronological test has high recall and low precision, because a performance gap might have other causes. Used together, one flags and the other confirms.

A worked version for anyone evaluating a new model this month. Take the datasets you care about, sort by release date, and mark the model cutoff. Report results for the two halves separately, and report how often the model beats the majority baseline in each half rather than mean accuracy. If the pre-cutoff half is much better, do not headline the aggregate number. If you want to go further, prompt the model to produce training examples for each task and count how many come back verbatim.

What we would change in our own evaluations

We have been treating the few-shot number as a property of the model. This paper argues it is a property of the pair of model and dataset, and that the pairing is mostly about dates. From now on, any few-shot claim we make will come with the cutoff and the release date next to it, and we will report the post-cutoff subset on its own. For the benchmarks we build ourselves, the release date is a feature, and it decays. A dataset is a clean few-shot test only until the next crawl includes it.

The open question is whether the same pattern holds for the current generation. All twelve models here predate the models people are excited about now, and the post-cutoff datasets are small. Someone should rerun the chronology on models with 2023 cutoffs, using datasets released in the last six months, before those datasets are absorbed too.

Sources

  1. Li and Flanigan, Task Contamination: Language Models May Not Be Few-Shot Anymore (arXiv 2312.16337)