The test-of-time pick

The 2023 test-of-time award went to Distributed Representations of Words and Phrases and their Compositionality, the second word2vec paper, presented at NeurIPS 2013 by Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado and Jeffrey Dean. The citation puts it above 40,000 references and credits it with beginning a new era in natural language processing. Ten years is the right distance to judge that, and the judgement is hard to argue with.

The author list is its own commentary on the decade. Sutskever went on to co-found OpenAI and to co-author the papers that defined the current era. Dean ran the effort at Google that produced much of the surrounding infrastructure. A 2013 paper about learning word vectors by predicting neighbours seeded a set of ideas and a set of people who then built what came after. The award is for the paper and reads, at ten years' distance, as being for the lineage too.

What the main track rewarded

Two papers won outstanding main-track awards, and the pairing is telling. One is Privacy Auditing with One (1) Training Run by Thomas Steinke, Milad Nasr and Matthew Jagielski, which audits a differentially private system with a single run and minimal assumptions in both black-box and white-box settings. The other is Are Emergent Abilities of Large Language Models a Mirage, by Rylan Schaeffer, Brando Miranda and Sanmi Koyejo, which argues that claimed emergent abilities can evaporate under a different metric or better statistics and may not be a fundamental property of scale.

That a skeptical paper about one of the year's most repeated claims took the top award is worth sitting with. The field spent 2023 talking about emergence as if it were settled, and the conference gave its award to the paper saying the phenomenon might be an artifact of how the metric was chosen. The runner-up list carried the constructive counterpart, Scaling Data-Constrained Language Models by Muennighoff and colleagues, which found that up to four epochs of repeated data barely change the loss, and Direct Preference Optimization, the DPO paper by Rafailov and colleagues, which showed you can skip fitting a separate reward model.

The datasets and benchmarks track

The datasets and benchmarks awards went to ClimSim, a climate emulation dataset with 5.7 billion input-output pairs led by Sungduk Yu and Walter Hannah with a large team, and to DecodingTrust, a trustworthiness assessment of GPT models across eight dimensions that found the models can be easily steered into toxic and biased output. DecodingTrust winning here is a small sign of where the venue's attention was going, since it is an evaluation of a commercial model that most attendees could not inspect, published at a conference that historically rewarded methods.

The scale of the event frames all of this. NeurIPS 2023 took 13,300 submissions and accepted 3,540, with 502 papers flagged for ethics review. Those are numbers for an industry, not a workshop, and the awards are being chosen from a pool large enough that the selection is a statement about priorities as much as a ranking.

The work that was not there

Here is what struck us reading the list. The models that defined 2023, GPT-4, Claude, the instruction-tuned systems people actually used, were not presented at NeurIPS as papers. GPT-4 arrived as a technical report with the architecture withheld. The DecodingTrust award is an evaluation of one of those systems from the outside, which is the closest the program comes to the frontier, and it comes as an audit rather than as the thing being audited.

So the awards tell a coherent story with a gap in the middle. The test-of-time award honours the academic paper that started the era. The main-track awards honour careful academic work, some of it questioning the hype. And the most consequential systems of the year are present only as objects other people study, because the labs that built them stopped publishing the details. The conference is thriving and the frontier has partly left the room.

What we take from it

The healthy reading is that this is a division of labour the field can live with. Companies ship systems, the academic venue does the auditing, the scaling analysis and the skeptical statistics, and a paper like the emergence one is exactly the contribution a university group is placed to make and a lab is not incentivised to. DecodingTrust and the mirage paper are both that kind of work, and both won.

The less healthy reading is that a venue whose test-of-time award goes to a paper whose lead authors left to build closed systems is honouring a pipeline that no longer feeds it. We do not know which reading is right. We do know that the 2013 winners could publish the method in full and the 2023 equivalents could not, and that is the change the award list documents without naming.

Sources

  1. NeurIPS blog, Announcing the NeurIPS 2023 Paper Awards (December 11, 2023)