Gemini 1.0 and the bet on native multimodality
Google's first Gemini models were trained on text, images, audio and video from the start rather than attaching a vision encoder to a finished language model. What the technical report says that phrase means, and what the early numbers do and do not show.
What was announced
Google announced Gemini on 6 December in three sizes. Ultra is the largest and is held back until early next year pending safety evaluation. Pro is the general-purpose tier and went into Bard immediately, with API access through AI Studio and Vertex from 13 December. Nano is for on-device use and comes in two sizes, 1.8 billion parameters for Nano-1 and 3.25 billion for Nano-2, shipped 4-bit quantised on the Pixel 8 Pro. The technical report followed on arXiv, credited to a Gemini Team of well over a thousand contributors.
The claim Google leads with is that Gemini Ultra beats the previous state of the art on 30 of 32 academic benchmarks and is the first model to exceed human expert performance on MMLU, at 90.0 percent. The number we find more interesting is buried in the report. That 90.04 percent was obtained with chain-of-thought prompting and 32 samples. With a standard 5-shot prompt the same model scores 83.7 percent. Both are strong results and only one of them is comparable to how MMLU has historically been reported.
What natively multimodal means in the report
The phrase in the announcement is that Gemini was built from the ground up to understand text, code, audio, images and video. The report is more specific. The models are Transformer decoders with a 32k context window and multi-query attention, trained on Google TPUs. Visual encoding is described as inspired by Google's own Flamingo, CoCa and PaLI work, with the stated distinction that these models are multimodal from the beginning and can natively output images using discrete image tokens.
Audio comes in as Universal Speech Model features at 16 kHz rather than being transcribed to text first, so the model sees something closer to the sound than to a transcript. Video is handled by encoding it as a sequence of frames within the long context window, which is a plain approach that relies on the 32k context rather than a dedicated video encoder. The pretraining data spans all of these modalities together, and that joint training from the start is what the word native is doing.
The contrast the report is drawing is with what we will call the late-fusion recipe. Train a language model on text. Separately train a vision encoder. Then connect them with a small adapter and a modest amount of paired data. That recipe has produced most of the open multimodal models of the past year, and it is cheap, because the expensive part, the language model, is reused. Gemini's bet is that a model which learned to read pixels and audio at the same time it learned to read text will understand them better than one that had vision bolted on afterwards.
What the numbers show so far
The strongest evidence for the bet is on the harder multimodal benchmarks. On MMMU, which tests college-level questions that require reading diagrams and images, Gemini Ultra scores 59.4 percent pass@1 and 62.4 percent with majority voting over 32 samples, which the report describes as more than five points over the previous best. Google also states that the image benchmark results were obtained without OCR assistance, meaning the model read the text in images itself rather than being handed a transcript.
The audio results are the ones we did not expect. On FLEURS across 62 languages, Gemini Pro reaches a 7.6 percent word error rate against 17.6 percent for Whisper. That is Pro, the middle tier, beating a dedicated speech model by a wide margin on a multilingual test. If native audio training is responsible for that gap, it is the clearest single data point in the report in favour of the approach.
The evidence is still short of a controlled comparison. Nothing in the report holds compute fixed and trains a late-fusion model of equal size on the same data, so we cannot separate the contribution of joint training from the contribution of scale, data and TPU budget. The benchmark gains are real. Whether they came from native multimodality or from being a very large model with a lot of image data is a question the report does not answer and probably was not designed to.
The case for and against the bet
The argument for joint training is that late fusion caps the depth of cross-modal understanding at whatever the adapter can carry. A vision encoder trained on captioning learns features that are good for captioning. A model that sees pixels during language pretraining can learn features that serve reasoning, because reasoning is what the training objective rewards. The MMMU result, where the questions require reading a chart and then doing several steps of inference, is the kind of task where that should show up first.
The argument against is cost and flexibility. Late fusion lets a lab upgrade the vision side without retraining the language side, and vice versa, and it lets an open-weights community build a multimodal model in a week from a released language model. Native training means every modality is baked into one expensive run. If the gains are modest, the field will keep bolting encoders on, and the report gives it no controlled reason not to.
We would like to see someone train two models at a few billion parameters, one with images in pretraining from token zero and one with a late adapter, matched on total compute, and run both on MMMU and FLEURS. That experiment is affordable outside Google and it would tell us whether the word native is doing real work. Until then the honest summary is that Gemini Ultra is a very strong model whose report makes a plausible architectural argument it does not yet test.
Sources
From the foundation