The Munk debate: Bengio and Tegmark versus LeCun and Mitchell
Four senior researchers argued whether AI research poses an existential threat in front of a Toronto audience that voted before and after. The vote moved three points. Notes on what the format could and could not settle.
The setup
On June 22 the Munk Debates put the resolution to a hall at Roy Thomson Hall in Toronto: be it resolved, AI research and development poses an existential threat. Max Tegmark and Yoshua Bengio argued for. Melanie Mitchell and Yann LeCun argued against. As far as we know this is the only time four researchers of this standing have argued the existential risk question in a formal adversarial format with a scored outcome, and that alone made it worth watching.
The scoring is what the Munk format is built around. The audience votes on the resolution before the debate and again after, and the side that moves more of the room wins. Before the debate 67 percent of the audience supported the resolution and 33 percent opposed it. After, it was 64 to 36. The con side won on a three point swing.
What the three points mean
Read one way, Mitchell and LeCun won. Read another, two thirds of a self selected Toronto audience walked in believing AI research is an existential threat and two thirds walked out believing the same thing. Three points is within the range you would get from people reconsidering a vaguely worded proposition on a second reading. We do not think either side changed many minds, and we doubt either side expected to.
The resolution wording did a lot of work. Poses an existential threat can be read as could plausibly lead to catastrophe if mishandled, which is a claim LeCun has at times been willing to grant, or as is on a path to doing so, which he flatly rejects. The pro side only needed the room to accept the first reading. The con side needed it to demand the second. A debate that turns on the reading of one verb is not a debate that produces a finding.
The arguments that were actually on the table
The Munk framing of the pro case listed the concerns in ascending order of severity: deep fakes and mass propaganda, the concentration of economic and political power in the corporations that control the systems, and the possibility that we lose control of powerful AIs pursuing objectives that harm humanity. That ordering matters. The first two are near term, observable, and hard to argue against. The third is the existential one, and it is the only one the resolution is about.
The con case, as framed, was that ChatGPT's arrival heralds a wave of productivity and creativity and that the risks come from rapid and unregulated development rather than from the technology itself. That is a claim about governance rather than about capability, and it is the line that lets a sceptic agree that there are serious dangers while voting against the resolution. Our guess is that most of the three point swing came from people who found that distinction persuasive on the night.
Why the format could not settle it
A Munk debate is a persuasion contest with a fixed clock, and the existential risk question is an argument about probabilities over decades where neither side can point to an experiment. Tegmark and Bengio could not produce a rogue system. Mitchell and LeCun could not produce a proof that scaling stops short of one. Both sides were left arguing from intuition about what future models will be like, and intuitions about that had already diverged in this field long before anyone bought a ticket.
The audience vote also measures the wrong thing. It measures how well each pair performed in front of a crowd, which selects for the debaters' rhetorical skill and the crowd's priors. A three point movement toward the con side tells us LeCun and Mitchell were slightly more effective at the podium. It tells us nothing about whether they are right, and treating it as evidence on the underlying question would be a category error that both sides, to their credit, avoided.
What we would want instead
The useful output of a debate like this would be a list of the concrete disagreements that survive two hours of cross examination. From the material we have seen, at least two survived. One is whether the current line of systems can acquire goals that are not written into them, which is an empirical question that evaluation groups are beginning to test. The other is whether regulation of development is sufficient on its own, which is a policy question that no debate will answer.
We would rather see the four of them asked to write down, in advance, what evidence would move each of them, and then reconvene when some of it exists. That would be a worse show and a better experiment. The Toronto vote showed that a good show does not move a room much on a question like this, so we may as well try the experiment.
Sources
From the foundation