What the paper claimed

On 22 March 2023, fourteen Microsoft Research authors led by Sébastien Bubeck posted a long report on an early version of GPT-4 and gave it a title nobody in the field has stopped arguing about. The abstract says the model "can solve novel and difficult tasks that span mathematics, coding, vision, medicine, law, psychology and more, without needing any special prompting," that its performance "is strikingly close to human-level performance," and that it "could reasonably be viewed as an early (yet still incomplete) version of an artificial general intelligence (AGI) system."

We read it the week it came out, like everyone else, and we remember the feeling of reading the transcripts. It was the first document that made a lot of working researchers, ourselves included, revise their sense of what the next few years would look like. Three years later we went back to it with a pen, and this note is what we found.

The paper is careful in places the title is not. The authors say explicitly that the version they tested was an early checkpoint, that they investigated limitations alongside capabilities, and that moving toward more general systems will need paradigms "beyond next-word prediction." Those caveats were in the text from day one. They just were not what anyone quoted.

What held up

The qualitative claim held up better than its critics expected. The paper said the model could handle unfamiliar tasks across many domains without task-specific prompting, and everything that has shipped since has made that description feel understated rather than inflated. If the argument in 2023 was whether a next-token predictor could do anything that deserved the word "general," that argument is over in practice, whatever one thinks of the word.

The paper also did something methodologically useful that we underrated at the time. Instead of running the standard benchmark suite, it built small, weird tasks designed so the answer could not have been memorized, and it showed the prompts. That style, probing with hand-built cases and publishing the transcripts, became the default way people investigated capabilities in the following two years. A lot of what we now call evals descends from that section of the paper more than from any leaderboard.

What could not be checked

Here is where the rereading gets uncomfortable. Gary Marcus published his response within days, under the title "The Sparks of AGI? Or the end of science?", and his central point was not about whether the model was impressive. It was that the paper disclosed "not the architecture, nor the training set. Nothing." His line that the core claim "literally cannot be tested with serious scrutiny, because the scientific community has no access to the training data" was true in March 2023 and it is still true. No one outside Microsoft and OpenAI has ever been able to rerun the paper.

Marcus also pointed at training data contamination. The paper argues its tasks were novel, but the argument rests on the authors' own belief about what was and was not in the corpus, and nobody can audit that belief. He noted that OpenAI had "begun to incorporate user experiments into the training corpus, killing the scientific community's ability to test the single most critical question: the ability of these models to generalize to new test cases." Once user prompts feed the next version, every clever probe someone posts is potentially in the next training set.

The most damaging item is in his postscript. Researchers who tried to rerun the paper's prompts against the publicly available, RLHF-tuned GPT-4 "failed to replicate 4/5 of their prompts," and Marcus described the results as "completely irreproducible." The paper's own explanation is that the tested checkpoint was earlier and less constrained than the public model. That may be true, and it is also exactly the kind of explanation that can never be falsified when the checkpoint is not available.

We want to be fair to the authors. A long qualitative report on a closed model was not pretending to be a controlled experiment, and it says so. But Marcus's harder charge was about the form. Companies were, in his words, "pretending to be contributors to science, formatting their work as science with graphs and tables and abstracts" while shipping "extraordinarily powerful yet unreliable systems." Rereading the paper now, with its abstract and figure numbering and appendix structure, we find it hard to say he was wrong about the form.

What the episode taught about evaluating closed models

The practical lesson we take from the three years is that a capability claim about a closed model has a shelf life of roughly one deployment. The thing that was tested is gone. The public model is a different artifact, tuned differently, and it will itself be replaced. If the evidence is a set of transcripts from an unavailable checkpoint, the evidence expires the moment the checkpoint does. We mean that as a description of the object, with no judgment of the authors attached.

The second lesson is that the "novel task" defence needs a paper trail. When we run a probe at the foundation now, we write down when the prompt was authored, where it has been posted, and which model versions have seen it, because the alternative is the situation Marcus described, where the field cannot tell generalization from recall. That is tedious bookkeeping and it is the only thing that makes a later replication attempt meaningful.

The third lesson is about who gets to name things. The word "sparks" did a great deal of work in 2023. It let the paper claim something enormous while retaining deniability about how much. Marcus's jab that by the same logic "a calculator" or "Eliza" or "Siri" might qualify was unkind and also identified the real problem, which is that the paper never committed to a definition that could fail. A claim that cannot fail is not a scientific claim, however good the transcripts are.

What we would want someone to do now

The obvious experiment has still not been run properly. Take the paper's published prompts, run them against every open-weight model released since with a documented training cutoff, and record which of the "novel" tasks are solved by models that could not have seen the paper and which are solved only by models that could. That would say something real about contamination, and open weights make it possible in a way GPT-4 never allowed.

The less obvious thing we would want is for the next paper of this kind, and there will be one, to ship the checkpoint or not ship the claim. Three years of hindsight has not changed our view that the Sparks paper was interesting. It has changed our view of what interesting is worth when nobody can check it.

Sources

  1. Sparks of Artificial General Intelligence: Early experiments with GPT-4 (arXiv)
  2. Gary Marcus, "The Sparks of AGI? Or the end of science?"