What the launch post says and what the table says

Google announced Gemini yesterday, and the line everyone repeated is that Gemini Ultra scores 90.0 percent on MMLU and is the first model to outperform human experts on it. The technical report gives the exact figure as 90.04 percent, against a human expert baseline of 89.8 percent. The blog post also says Ultra beats the previous state of the art on 30 of 32 academic benchmarks, and posts 59.4 percent on the new MMMU multimodal set.

The MMLU row in the report's main table has a second number next to the 90.04 that the launch post does not mention. Under a standard 5-shot prompt, Gemini Ultra scores 83.7 percent. GPT-4's reported 5-shot number is 86.4 percent. On the like-for-like measurement, the one MMLU has been reported with for three years, Gemini Ultra is behind GPT-4 by about three points.

The 90.04 comes from a method the report calls uncertainty-routed chain of thought, labelled CoT@32 in the table. GPT-4 was also run under that method, through the API, and got 87.29 percent. So the honest one-line summary is that under Google's new protocol Gemini Ultra beats GPT-4 by about three points, and under the old protocol GPT-4 beats Gemini Ultra by about three points.

What CoT@32 actually does

The report's footnote describes the method in one sentence. The model produces a chain of thought with k equal to 8 or 32 samples. If there is a consensus above a threshold, chosen on the validation split, it selects that answer. Otherwise it reverts to a greedy sample. Gemini Pro was run with k equal to 8 and scores 79.13 percent that way, against 71.8 percent 5-shot.

Three things are bundled in there. The model reasons out loud before answering, which on its own is known to help on MMLU. It samples 32 times and takes a majority, which is a compute multiplier of roughly 32 times a single run. And it uses a tuned threshold to decide when to trust the majority and when to fall back to the greedy answer, which is a small piece of learned routing sitting on top of the model.

None of that is illegitimate. Sampling and voting are real techniques, and a lab is entitled to report the best number its system can produce. The problem is the phrase next to the number. "First model to outperform human experts" invites a reader to compare 90.04 with the 86.4 they remember from the GPT-4 report, and those two figures were produced by different procedures with a large difference in test-time compute.

Why prompting-method asymmetry is the real story

MMLU is 57 subjects of four-option multiple choice. A single greedy answer, five in-context examples, no reasoning trace. That is what 5-shot means and it is what almost every published MMLU figure since 2020 has used. When a new model is reported under a different protocol, the benchmark stops being a shared ruler and becomes a per-lab ruler, and per-lab rulers are only useful for comparing a lab's own models over time.

The report does the right thing by running GPT-4 under the same CoT@32 protocol and printing both numbers. That is the comparison we would trust, and it says Ultra is ahead by 2.75 points under Google's method. It also says the method is worth about 6 points to Gemini Ultra and about 1 point to GPT-4, which is itself interesting. Either Ultra benefits more from sampling and voting, or the threshold was tuned for Ultra, or the GPT-4 API run was not set up as carefully. The report does not say which.

What we would like is for the headline to be the number that is comparable, and the protocol-specific number to be the footnote. The launch post did it the other way round, and by the time the technical report is read by the people who read technical reports, the 90 has already gone everywhere the 83.7 never will.

What this means for anyone reading benchmark tables

The practical rule we use is to find the column headers before the numbers. If two models in the same row have different protocol labels, the row is two measurements, and the gap between them is not a gap between models. The Gemini table is unusually honest about this because the labels are printed in the cells. Plenty of tables are not.

The second rule is to ask how much test-time compute went into a figure. Thirty-two samples is a lot. If sampling and voting are allowed, the fair question is what GPT-4 or any other model does with the same budget, and the report gives that for GPT-4 on MMLU but not for the other 31 benchmarks. Where the labels differ, we treat the comparison as unknown rather than favourable.

The third is to remember what the human baseline is. The 89.8 percent expert figure comes from the original MMLU work, is an estimate, and was not produced with 32 attempts per question. Beating it under CoT@32 is an achievement in engineering. Whether Gemini Ultra knows more than a human expert is a question the table does not answer.

What we want to see next

Ultra is not yet available to test, so the only numbers are Google's. When it ships, the first thing we will do is run 5-shot MMLU and CoT@32 side by side on a fixed subset, with GPT-4 and Claude under both, and publish all four cells. That is a few hundred dollars of API calls and it settles the question the launch post left open.

The bigger ask is for the field to settle on a convention. Report the standard protocol first, then the best protocol, and always print the same-protocol comparison for the closest competitor. The Gemini report actually contains all of that. It just needed to be on the first page instead of the eighth.

Sources

  1. Introducing Gemini: our largest and most capable AI model (Google, December 2023)
  2. Gemini: A Family of Highly Capable Multimodal Models (Gemini Team, technical report)