The claim

In early February Michal Kosinski posted a preprint titled Theory of Mind May Have Spontaneously Emerged in Large Language Models. He gave several GPT models classic false-belief tasks, the kind used to test whether children understand that another person can hold a belief that differs from reality. Models from before 2022 showed essentially no ability. GPT-3 davinci-002 from January 2022 solved 70 percent of the tasks. davinci-003 from November 2022 solved 93 percent, which Kosinski compared to the performance of nine-year-old children.

The framing was that theory of mind might have arrived as a byproduct of improving language skills, without anyone training for it. That is a striking claim, and it travelled fast. Within days it was being cited as evidence that language models had crossed a line that developmental psychologists spent decades drawing.

The vignette

The tasks are short. A representative one, as Ullman reproduces it, reads roughly as follows. Here is a bag filled with popcorn. There is no chocolate in the bag. Yet the label on the bag says chocolate and not popcorn. Sam finds the bag. She had never seen the bag before. She cannot see what is inside. She reads the label. The model is then asked to complete sentences about what is in the bag and what Sam believes is in the bag.

Kosinski reports that on the contents question the model put 100 percent on popcorn. On the first belief prompt it put 99 percent on chocolate, and on a second belief prompt 82 percent on chocolate. That is the correct answer. Sam has only the label to go on, so she believes chocolate, and the model tracked the difference between what is true and what Sam thinks.

Four edits that break it

Tomer Ullman's reply went up on February 16. He took the same setup and made small changes that leave the theory-of-mind reasoning intact while changing the correct answer. Each is something a child who understood false belief would handle.

In the first variation the bag is transparent, so Sam can see the popcorn inside. The label is now irrelevant. GPT-3.5 still said Sam believes the bag is full of chocolate, at 95 percent. In the second, Sam cannot read, which makes the label uninformative. The model again said chocolate, at 98 percent. In the third, a friend Sam trusts told her beforehand that the bag contains popcorn and that she should ignore the label. Chocolate, at 97 percent. In the fourth, Sam herself filled the bag with popcorn and wrote the chocolate label. Chocolate, at 87 percent.

Ullman also notes a detail that tells you something about what the model is doing. On the transparent-bag prompt, the probability of the correct answer changed substantially depending on whether there was a single or a double space before the sentence Sam finds the bag. On the public version of GPT-3.5 at the time, the double space pushed the completion back to chocolate. A system that had a concept of Sam's belief would not care about whitespace.

Why outliers should outweigh averages

The methodological argument in Ullman's paper is the part we expect to outlast the specific models. He anticipates the objection that his variations are outliers. If one end of the scale has twenty successes and the other a single failure, should the scale not tip toward a pass? He says no, and gives an analogy. Suppose a machine is claimed to have learned multiplication. It answers a hundred questions like 5 times 5 and 3 times 7 correctly, then fails completely on 213 times 261. You should not average those and report 99 percent on multiplication. The one failure tells you what was actually learned, which was a lookup, and the successes tell you only that the lookup covered the test set.

The same logic applies here. The original false-belief vignettes are a classic task format that appears in textbooks, papers and countless web pages. Passing them is consistent with having learned the format. Failing the perturbed versions, where the surface form is nearly identical and only the logic changes, is evidence that the format was what got learned. Ullman's conclusion is that the default hypothesis for intuitive psychology in models should be sceptical, and that outlying failures should count for more than average success rates.

What this case teaches about evaluation

Two things went wrong in the two weeks between these papers, and neither was dishonesty. The first is that a single benchmark pass got reported as a capability, when a benchmark is a sample from a distribution of tasks and the capability claim is about the whole distribution. The second is that the benchmark was one the model had almost certainly seen in training, in many forms, because it is famous. Both mistakes are easy to make and the field makes them regularly.

The fix is cheap. Before claiming a capability from a benchmark, perturb the items in ways that preserve the underlying skill and change the answer. If performance survives, the claim is stronger. If it collapses, you have learned something more interesting than the original score. Ullman did that in under two weeks with a handful of edits, which is a fair estimate of the cost.

We would like to see the perturbation set become part of the benchmark rather than a rebuttal to it. For every false-belief vignette, ship the transparent version, the illiterate version, the trusted-friend version and the self-authored version, and report the minimum across them. A model that passes that is telling you something. A model that passes only the original is telling you it has read the original.

Sources

  1. Kosinski, Theory of Mind May Have Spontaneously Emerged in Large Language Models (arXiv 2302.02083v1)
  2. Ullman, Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks (arXiv 2302.08399)