Lost in the middle: the U-shaped curve every RAG pipeline had to learn
Liu et al. moved the answer-bearing document around a 20-document context and watched GPT-3.5-Turbo go from 75.8 percent to 53.8 percent, below its own closed-book score. Reading notes on the paper that changed how everyone orders their retrieved chunks.
The experiment
Nelson Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni and Percy Liang built a multi-document question answering task where exactly one of the supplied documents contains the answer. They then moved that document to different positions and changed nothing else. Same question, same set of documents, same model. The only variable is where the useful text sits.
With 20 documents, GPT-3.5-Turbo answers 75.8 percent correctly when the answer is in the first document, 53.8 percent when it is ninth, and 63.2 percent when it is last. The comparison that should stop you is the closed-book number. Given no documents at all, the same model scores 56.1 percent. Put the answer in the middle of a context and the model does worse than if you had not retrieved anything.
They ran the same shape of test on MPT-30B-Instruct, LongChat-13B with 16K context, GPT-3.5-Turbo with 16K, and Claude-1.3 with both standard and 100K contexts. The U shape shows up broadly. Both MPT-30B and MPT-30B-Instruct show it, so this is not an artifact of instruction tuning.
A second task with no semantics at all
The multi-document result could be about reasoning, so they added a synthetic key-value retrieval task. Give the model a JSON object of key-value pairs, give it a key, ask for the value. There is nothing to understand, only something to find. With 300 pairs, GPT-3.5-Turbo (16K) degrades badly in the middle positions while Claude-1.3 and Claude-1.3 (100K) are near perfect throughout.
That is a useful pair of facts. The positional weakness is not universal across models, which means it is a property of training rather than a law of transformers. And it appears in a task with no reasoning in it, which means it is about locating information rather than about using it.
One intervention helped a lot. Query-aware contextualization, which places the question both before and after the documents rather than only after, brought key-value retrieval to near perfect, and to perfect for GPT-3.5-Turbo (16K). On the multi-document task it helped much less, which is worth remembering before treating it as a general fix.
What this does to a retrieval pipeline
The immediate consequence is about ordering. If your retriever returns 20 chunks sorted by score and you paste them in that order, the second-ranked chunk lands where the model reads worst. Ordering so that the highest-scoring chunks occupy the first and last positions costs nothing and follows directly from the curve.
The second consequence is about how many chunks to send. The paper runs an open-domain QA case study on Wikipedia and finds that model performance saturates long before retriever performance does. Going from 20 retrieved documents to 50 improves accuracy by about 1.5 percent for GPT-3.5-Turbo and about 1 percent for Claude-1.3, while the retriever is still finding more correct documents in that range. You are paying for tokens that the reader cannot use.
The third is about reranking. A reranker was previously a way to improve which chunks you send. This result makes it a way to decide where each chunk goes, which is a different and cheaper job, and it raises the value of being precise about the top few rather than broadly good across the top fifty.
What we would be careful about
This is a measurement of six models in mid-2023, not a theorem. Claude-1.3 was already flat on key-value retrieval, which suggests the curve can be trained away, and we expect the next generation of long-context models to be evaluated on exactly this and to look better. Position-sensitivity numbers age faster than most benchmark numbers, so anyone building on this should rerun the probe on their own model rather than citing the 2023 figures.
The durable contribution is the protocol. Take a task, hold the content fixed, vary only position, and plot accuracy against position. That takes an afternoon, works on any model with any context length, and gives you a curve you can point at when someone proposes to stuff a 100,000 token prompt. The paper says its aim is to provide new evaluation protocols for future long-context models, and that is the part we would keep.
The claim we would resist is that a long context window replaces retrieval. What this shows is that a window is a capacity, and using the middle of it is a capability that has to be demonstrated per model. Until someone shows us a flat curve on their own model, we will keep sending fewer chunks and putting the best ones at the edges.
Sources
From the foundation