The needle in a haystack test: one engineer's chart becomes an industry benchmark
Greg Kamradt hid one sentence in a stack of Paul Graham essays and plotted where GPT-4 and Claude 2.1 could find it. The heatmap is now the standard picture of long-context recall. Here is what it measures and what it does not.
The setup
The test is small enough to describe in a paragraph. Take a large body of unrelated text, in the original runs a collection of Paul Graham essays, and insert one sentence the model could not know from training. The needle Kamradt used reads: the best thing to do in San Francisco is eat a sandwich and sit in Dolores Park on a sunny day. Then ask the model what the best thing to do in San Francisco is, and check whether the answer contains the sandwich.
The two variables are context length and needle depth. You sweep the total length of the prompt from a few thousand tokens up to the model's limit, and for each length you place the needle at a range of depths from the top of the document to the bottom. Each cell in the resulting grid is a pass or a fail, and the whole thing is coloured as a heatmap. GPT-4 with the 128K window was run on November 8 and Claude 2.1 with its 200K window on November 21, and the repository still keeps those original results.
Why the picture travelled
Until this month the long-context claims from model providers came as a single number, the size of the window. A window size tells you what the tokenizer and the serving stack will accept. It tells you nothing about whether the model attends to token 90,000 as well as token 900. Kamradt's grid answered the second question in one image, with length on one axis and depth on the other, and it did so cheaply enough that anyone with an API key could rerun it.
That is why every long-context release since has shipped with its own version. The format is easy to read, hard to argue with on its own terms, and embarrassing when it goes red. It is also easy to game, which we will come back to.
The Claude 2.1 episode
The most instructive thing that happened with the test happened after the chart was published. Anthropic reproduced the setup internally and reported that Claude 2.1 initially found the sentence in only 27 percent of cases, frequently answering that the document did not contain enough information. Their fix was a single line added to the start of the assistant's turn: here is the most relevant sentence in the context. With that prefix in place the score went to 98 percent.
Anthropic's explanation was that the model had been trained with a lot of feedback on real document tasks aimed at reducing unsupported claims, and a sentence about sandwiches dropped into an essay about startups looks out of place. The model treated the needle as noise and declined to cite it. We find this more interesting than the original chart. A retrieval failure in the heatmap turned out to be a refusal policy interacting with an artificial test, and a prompt change flipped it. Any benchmark whose score moves by 71 points on one line of prompt is measuring the prompt as much as the model.
What passing does not tell you
A green grid means the model can locate one salient, out-of-place sentence and repeat it. That is retrieval of a single fact with no competing candidates. It is a necessary condition for anything useful at long context, and we are glad we can now check it in an afternoon. But it is nowhere near sufficient.
Real documents do not contain one anomalous sentence. They contain many related facts, some of which contradict each other, and the question you care about usually needs two or three of them combined. The test says nothing about whether a model can count occurrences, notice that page 40 disagrees with page 900, or reason over a chain of references. The repository has since added variants that spread multiple needles through the context and a chain task where the model must follow links from A to B to C without being told a chain exists. Those are steps in the right direction, and we would trust a multi-needle or chain result far more than the single-needle grid that gets reproduced in launch posts.
The other caveat is contamination of the setup itself. The sandwich sentence and the Paul Graham corpus are now public and widely copied. A model trained after this month may have seen the exact pairing. Anyone rerunning the test should swap the needle and the haystack, which the tool supports.
What we would want next
We would like to see the heatmap kept as a smoke test and demoted from headline. Ship it, but ship it next to a multi-needle version, a version with distractor sentences that almost match, and a task that requires combining facts from two depths. Report the prompt template used, since the Claude 2.1 case shows it can be the whole result.
And we would like someone to run the grid with a needle that fits the document, a plausible sentence about startups buried in essays about startups, rather than one that stands out. If the model finds a sentence because it looks strange, we are measuring anomaly detection, which is a fine thing to measure but is not what people mean when they say a model can read a book.
Sources
From the foundation