The claim

A paper posted this week by Rylan Schaeffer, Brando Miranda and Sanmi Koyejo argues that emergent abilities in large language models are an artefact of how we score them. The abstract puts it bluntly. Emergent abilities appear because of the researcher's choice of metric rather than because of fundamental changes in model behaviour with scale. The authors say that nonlinear or discontinuous metrics produce apparently sharp jumps, and that linear or continuous metrics applied to the same outputs show smooth, predictable improvement.

We have read the paper twice now and we think it splits into two claims of very different strength. The first is a statistical observation about scoring rules, and it is correct. The second is an interpretive claim that emergent abilities may not be a fundamental property of scaling at all. That one is contested, and the evidence in the paper supports it less than the framing suggests.

What the arithmetic experiment shows

The cleanest demonstration uses integer arithmetic on the InstructGPT and GPT-3 family, chosen because those models are publicly queryable. Scored by accuracy, which requires every token of the answer to be correct, the smaller models sit at zero and the largest model jumps up. Scored by token edit distance, which gives partial credit per token, the same outputs improve smoothly with scale. Nothing about the models changed between the two plots. Only the scoring rule did.

The authors then make a prediction that follows from this. If per-token error decays gradually with scale, then accuracy on a target string of length N should fall roughly geometrically with N, and small models should show above-chance accuracy rather than zero. They tested this by generating more test data to raise the resolution of the measurement, and found that every model in the family scores above chance on both arithmetic tasks. The zeros in the original plots were a resolution limit, set by one over the test set size, and not a property of the models.

This part of the argument has three named ingredients. A metric that scales the per-token error rate nonlinearly or discontinuously. Too few test examples to resolve small-model performance. Too few models sampled at the large end. Any one of them can manufacture a sharp corner in a curve that is smooth underneath.

The BIG-Bench audit

The second study is a meta-analysis of BIG-Bench, using hand-annotated claims of emergence. Emergent abilities appear under at most 5 of the 39 BIG-Bench metrics, and more than 92 percent of the claimed cases fall under just two of them, Multiple Choice Grade and Exact String Match. Multiple Choice Grade is a step function. Exact String Match is nonlinear in the length of the target. On tasks where LaMDA shows emergence under Multiple Choice Grade, switching to Brier Score, a proper scoring rule that BIG-Bench already reports, makes the emergence disappear.

The third study is the one people will remember. The authors induce emergent abilities on purpose in vision. Shallow autoencoders on CIFAR100 and small transformers classifying Omniglot characters show sharp jumps once you score them with a subset accuracy that demands every item in a sequence be right. The curves qualitatively match the published language model plots. If you can create emergence in a one-hidden-layer autoencoder by picking a metric, the metric is doing a lot of work.

Where the stronger claim overreaches

So far we agree with everything. Here is where we part company. The paper says in its related work that metric choice is likely wholly responsible for emergent abilities, and the discussion says such abilities may be creations of the researcher's choices. That is a claim about every reported emergence, and the evidence covers a specific subset of them, the ones where a smooth per-token metric is available and the task decomposes into tokens that can be partially correct.

Not every ability decomposes like that. Consider a task whose answer is a single yes or no, where there is no partial credit to award and the only continuous quantity is the model's probability on the right token. The paper's own recommendation would be to plot that probability. But a probability that rises smoothly from 0.5 to 0.9 across three orders of magnitude of compute can still correspond to a system that goes from useless to useful at some threshold in deployment. Smoothness in the underlying quantity does not make the practical transition any less abrupt, and it does not tell you in advance where the threshold sits.

The paper also acknowledges competing accounts without refuting them. Caballero and colleagues explain emergence as a change in the governing power law, under which the abilities are real. Michaud and colleagues argue emergence may be real under strong assumptions about the data. Schaeffer and colleagues show that emergence can be induced under a single power law, which is a different statement from showing that all observed emergence is induced that way.

What we take from it

The lasting contribution is methodological. Anyone who reports a sharp jump now has to answer two questions. Did you score with a continuous metric as well, and did you have enough test data to resolve the small models? If the jump survives both, it is interesting. If it does not, you had a scoring artefact. That is a real improvement in how the field argues, and we expect the paper to be cited mostly for it.

What we would like to see next is the same treatment applied to the abilities that matter for planning, where the smooth metric is not obvious. Chain-of-thought accuracy on multi-step problems, tool use that either completes or fails, and refusal behaviour under adversarial prompts do not have a token edit distance. If someone can find the continuous quantity underneath those and show it rising smoothly, the mirage claim gets much stronger. Until then we read this paper as a correction to plots, and a caution about the word emergent, rather than as proof that scale never surprises us.

Sources

  1. Schaeffer, Miranda and Koyejo, Are Emergent Abilities of Large Language Models a Mirage? (arXiv 2304.15004)