The chart that carried the launch

Google announced Gemini 1.5 Pro on February 15 with a standard 128,000 token context window and a limited preview of one million tokens through AI Studio and Vertex AI. The number that travelled was the needle-in-a-haystack result. In the announcement's words, the model found a specific piece of embedded text 99 percent of the time across blocks of data as long as one million tokens. Google said it had tested up to 10 million tokens in research, and the technical report puts the text figure at above 99.7 percent recall at one million tokens and 99.2 percent at 10 million.

For context, the announcement describes one million tokens as an hour of video, eleven hours of audio, codebases of more than 30,000 lines, or more than 700,000 words. The comparison points the report names are Claude 3.0 at 200,000 tokens and GPT-4 Turbo at 128,000. The architecture is mixture-of-experts, and Google claims performance broadly similar to Gemini 1.0 Ultra at lower compute, with 1.5 Pro beating 1.0 Pro on 87 percent of its evaluations.

What a needle test measures

The needle-in-a-haystack setup is simple. You take a long body of unrelated text, insert a single sentence with a fact the model would not otherwise know, and ask the model for that fact. You vary the total length and the depth at which the sentence is buried, and you plot the grid. A green grid at one million tokens means that attention, or whatever mechanism the model uses, can locate one salient sentence anywhere in a very long input.

That is a real result and we do not want to talk it down. Six months ago the practical ceiling for most people was 32,000 tokens and retrieval at the far end of the window degraded badly. But the test is a recall test with one target. It says nothing about whether the model can combine two facts from different places, count occurrences, or notice that a fact stated on page 40 contradicts one on page 900. The report authors say as much. They describe the haystack task as a retrieval task measuring recall and argue for evaluations that demand reasoning over multiple pieces of information scattered across a long context.

The number lower down the report

The same report contains a multi-needle variant that got far less attention. With 100 distinct needles hidden in the context and the model asked to retrieve all of them, Gemini 1.5 Pro holds about 70 percent recall up to 128,000 tokens and above 60 percent up to one million. GPT-4 Turbo, which cannot go past 128,000, averages around 50 percent at that length. So the model is still ahead of the alternative, and 60 percent at a million tokens is a good number for a task nobody could attempt a year ago.

But it is a very different picture from 99 percent. Going from one needle to a hundred drops recall by roughly 40 points at the longest length. If your application is summarising a long deposition, or extracting every date from a contract, or pulling every function that touches a particular table out of a codebase, the second chart is the one that describes your job. One-needle recall is the ceiling on the best case. Multi-needle recall is closer to the floor of the typical case.

How we would read a haystack chart now

The first thing we check is how many needles. A single needle answers the question of whether the window works at all. The second thing is whether the needle is lexically distinct from the haystack. In the original setups the inserted sentence is about something the surrounding text never mentions, which makes it an easy target for attention. A needle written in the same register as its surroundings, on the same topic, is a harder and more realistic test, and we have not seen a public number for that yet on this model.

The third thing is what happens after retrieval. The Kalamang result in the announcement is more informative than the haystack for this reason. The model was given a grammar manual for a language with very few speakers and learned to translate at a level Google describes as similar to a person who learned from the same materials. That requires reading, holding, and applying rules across a long document, which is the actual use case for a million tokens. We would like to see that kind of task turned into a repeatable benchmark rather than a demo.

What to try before trusting the window

If you have preview access, the experiment we would run first is a counting task. Put a known number of instances of a fact into a document, at random depths, and ask for the count and the locations. Then repeat with the instances phrased differently each time. That gives you a multi-needle curve for your own data and will tell you far more than the launch chart about whether you can replace a retrieval pipeline with a big prompt.

Our guess is that for a while the right pattern is both. Use the long window when the question needs the whole document held at once, and keep retrieval for anything where you need every instance found. The one-needle chart proves the window is real. The hundred-needle chart is the one that should set expectations, and we expect more variants along those lines, with reasoning steps rather than lookups, within the year.

Sources

  1. Google: Our next-generation model, Gemini 1.5 (February 15, 2024)
  2. Gemini 1.5 technical report (arXiv 2403.05530)