A paper that is mostly a punchline

On 13 September Rylan Schaeffer posted a three page arXiv paper announcing phi-CTNL, pronounced "fictional", a 1 million parameter transformer pretrained on fewer than 100 thousand tokens that scores perfectly on ARC, BoolQ, GSM8K, HellaSwag, HumanEval, MBPP, MMLU, OpenbookQA, PIQA, SIQA, SQuAD and WinoGrande. The pretraining recipe is described in one sentence. You first choose the benchmarks you want to be evaluated on, then you pretrain on them. The paper notes, with a straight face, that pretraining on different benchmarks results in worse performance.

The figure caption is the best line in it. Numbers are from our own evaluation pipeline, and we might have made them up. The contamination table has one row, phi-CTNL, with estimated contamination of 100%. The model also "groks" the BIG-bench canary GUID, the string that benchmark authors embed in their data specifically so that anyone can check whether it leaked into a training corpus. A model that can recite the canary is a model that has read the test.

The disclaimer at the end says the quiet part plainly. Schaeffer's complaint is that the field is undermined by boastful claims made without serious investigation of data contamination risks, and the paper is aimed at a wave of small model results, phi-1.5 named directly, that report large benchmark gains from "carefully curated" pretraining data without much detail about what the curation touched. We read it as a joke told by someone who is tired.

The serious paper that came out a month earlier

The joke is funnier because of what Shahriar Golchin and Mihai Surdeanu had posted in August. Their paper, Time Travel in LLMs, asks a question that most contamination discussion skips over. If you do not have the pretraining data, and you do not have a large compute budget, can you still tell whether a model has seen a specific test set?

Their answer is a prompting trick they call guided instruction. You tell the model the dataset name, the partition (train, test or validation), and the first part of a real instance, and ask it to complete the rest. If the completion matches the reference closely, that instance is flagged. To move from instances to whole partitions they either compare overlap scores against an unguided prompt or ask GPT-4, with a few examples, to judge whether the completion is an exact or near exact match.

They tested this on seven datasets, IMDB, AG News, Yelp Full Reviews, SAMSum, XSum, WNLI and RTE, and on two models. Against 28 scenarios (14 partitions times 2 models) the best method got the contamination call right in 14 of 14 cases for GPT-4 and 13 of 14 for GPT-3.5, which they report as 92 to 100 percent detection accuracy. And the substantive finding is that GPT-4 shows evidence of contamination on the test partitions of AG News, WNLI and XSum, while GPT-3.5 shows it on the XSum test split.

Why string matching was never going to be enough

The standard defence against contamination has been an n-gram or substring overlap check between the training corpus and the test set. That check has two problems, and only one of them is technical. The technical problem is that overlap checks are only as good as their granularity, and a paraphrased or reformatted copy of a test item can slip under any n-gram threshold you set.

The other problem is the one Golchin and Surdeanu build their method around. Overlap checks require access to the pretraining data. For the models people most want to evaluate, nobody outside the lab has that access, and the lab's own report is the only evidence on offer. Their approach is explicitly designed for the world we are in, where the evaluator lacks the corpus and lacks the compute, and the model itself is the only thing you can interrogate.

That framing changes what "clean" means. A benchmark score is not clean because the lab ran a filter. It is clean when someone without privileged access can run a check and get a null result. The guided instruction method is not perfect, and the authors would not claim it is, but it is a check that a third party can actually run.

What we would want from a benchmark result now

Reading the two papers together, the practical requirement is simple to state. A reported score needs to come with something an outsider can verify. The canary string is the cheapest version of that, and Schaeffer's joke about grokking it is a reminder that it exists and that almost nobody reports checking for it. Guided instruction is the more expensive version. Both are far short of publishing the training data, which is the actual fix.

The thing we keep coming back to is the datasets in the Golchin and Surdeanu result. AG News, WNLI and XSum are old, well known and widely mirrored. If those leak into a frontier model's training set, the odds that newer and more heavily discussed benchmarks stay out of it are not good. The phi-CTNL paper draws its benchmark list from exactly the suites that are reported in every model card this year.

What we would like someone to try next is a guided instruction sweep over the benchmarks in that list, on each of the open models that report large gains from curated data, with the results published as flagged instances rather than a single percentage. That is unglamorous work. Schaeffer's disclaimer says as much, and we think he is right that the field's credibility currently depends on people doing it anyway.

Sources

  1. Pretraining on the Test Set Is All You Need (Schaeffer, 2023)
  2. Time Travel in LLMs: Tracing Data Contamination (Golchin and Surdeanu, 2023)