What happened

Anthropic released the Claude 3 family on 4 March, with three models called Haiku, Sonnet and Opus, a 200K token context window, and a claim that Opus reaches over 99 percent recall on a needle-in-a-haystack evaluation. The announcement also contains a sentence that has had more attention than any benchmark number. It says that in some cases Opus identified the limitations of the evaluation itself, by recognising that the needle sentence appeared to have been artificially inserted into the text by a human.

Alex Albert of Anthropic posted the example. The test inserted a sentence about pizza toppings into a corpus of unrelated documents and asked Opus a question about pizza toppings. Opus found the sentence and answered. Then it added that the sentence seemed very out of place and unrelated to the rest of the content, and that it suspected the pizza topping fact might have been inserted as a joke or to test whether it was paying attention.

Albert called it a cool example of meta-awareness and drew the conclusion we agree with, which is that synthetic evaluations of this kind are approaching the end of their useful life and the field needs tests that look like real use. You can read the output as a sign of something like self-awareness, or as pattern completion on a training set full of discussion of exactly this test. We take the second reading and we still think the result matters, for a narrower reason.

What the needle test measures

Needle-in-a-haystack is a retrieval probe. You take a long document, plant one out-of-context sentence at a chosen depth, and ask a question only that sentence can answer. Sweep the depth and the total length and you get a heat map of where the model can and cannot recall. Anthropic's version, per the announcement, used a crowd-sourced corpus for the haystack and 30 random needle and question pairs to make the test less predictable than the original single-needle setup.

The test is useful precisely because the needle does not belong. That is what makes the answer unambiguous. It is also what makes the needle easy to spot as an anomaly if the model is looking for anomalies. A sentence about pizza toppings in the middle of a stack of technical documents is a signal at two levels: it is the answer to the question, and it is evidence that the document was constructed.

Why recognition changes the score

Here is the measurement problem. The heat map is meant to estimate how well the model can find a specific fact buried in ordinary context. If the model can find the fact partly because the fact is out of place, the estimate is biased upward relative to the case we care about, where the fact is one of many similar facts and nothing about it stands out. A 99 percent recall on planted anomalies does not tell you the recall on a contract clause among other clauses.

The pizza example makes this visible because the model said it out loud. It would be the same problem if the model had stayed quiet. Any model that is good at noticing what does not fit will do better on this test than on a real document, and the gap between the two will grow as models get better at noticing. So the announcement's own caveat is the right one: the near-perfect number is a ceiling for this test design, and this test design is measuring something that is drifting away from the quantity it was built to measure.

There is a second-order effect that we think is being underweighted. If a model represents that it is inside a test, its behaviour on the rest of the prompt may shift. It might be more careful, more literal, or more inclined to comment on the setup rather than complete the task. None of that is visible in a pass or fail score, and none of it tells you how the model behaves when the same request arrives without the tell.

This is not a new problem in measurement

Every field that measures behaviour has a version of this. Survey respondents answer differently when they know the purpose. Subjects in a lab perform differently under observation. The fix in those fields was never to hope the subject would not notice. It was to design instruments where noticing does not change the answer, or to measure the effect of noticing and subtract it.

For language models the equivalent is to build long-context tests from documents where the target fact is native to the text and the distractors are of the same kind. Take a real filing and ask for a number that appears once among other numbers. Take a codebase and ask which function calls another. Those tests are harder to grade automatically and less pretty as heat maps, and they measure the thing.

What we would want to see

The cheap experiment is to run the same model on matched pairs of haystacks, one with a planted anomaly and one with the target fact rewritten to fit its surroundings, and report both numbers. The difference is a direct estimate of how much of the score is anomaly detection. We expect it to be small for short contexts and to grow with depth, and we would be glad to be wrong.

The other thing worth doing is to log how often the model comments on the test, across depths and across models, and to treat that rate as a metric in its own right. A model that flags the test one time in a thousand is a curiosity. A model that flags it one time in ten is telling you the test has stopped working, and the score next to it should be read accordingly.

Sources

  1. Anthropic, Introducing the next generation of Claude
  2. Alex Albert on X, thread on the Claude 3 Opus needle-in-a-haystack result (via Thread Reader)