The problem with a chunk on its own

Retrieval-augmented generation usually starts by cutting documents into chunks of a few hundred tokens, embedding each one, and pulling the nearest chunks at query time. The failure this creates is that a chunk stripped from its document loses the information that made it findable. A sentence saying revenue grew 3 percent over the previous quarter does not say which company or which quarter, so neither an embedding nor a keyword index can connect it to a query that asks about that company.

Anthropic's note from September 19 proposes the simplest fix we have seen written up properly. Before embedding, run each chunk through Claude 3 Haiku with the whole document in context and a short prompt asking for a succinct sentence or two situating the chunk within the document for search. Prepend that text to the chunk. Then index the result both as an embedding and in BM25. They call the two halves contextual embeddings and contextual BM25.

The numbers

The metric is retrieval failure rate at top 20, the fraction of queries where the relevant chunk was not among the 20 retrieved. Their baseline is 5.7 percent. Contextual embeddings on their own bring that to 3.7 percent, a 35 percent reduction. Adding contextual BM25 brings it to 2.9 percent, a 49 percent reduction. Adding a reranker on top, with Cohere's reranker scoring the top 150 candidates and passing 20 to the model, brings it to 1.9 percent, a 67 percent reduction.

Two smaller details in the note matter as much as the headline. Retrieving 20 chunks beat retrieving 5 or 10, which is worth remembering when people argue for tight context windows on cost grounds. And the cost of generating context for every chunk, with prompt caching so the document is only paid for once, is about 1.02 dollars per million document tokens. That is cheap enough that the decision to do it is not really a cost decision.

Why BM25 is back in the pipeline

The part of this note that we think will age best is the explicit return to lexical search. BM25 is a term-frequency ranking function from the 1990s, and for a while the fashion was to treat embeddings as a strict replacement for it. The note's argument for keeping both is concrete. Embeddings capture meaning and miss exact strings. A query for an error code, a function name or a ticket identifier needs the exact string, and a dense vector will happily return something semantically nearby and useless. BM25 gets those right and misses paraphrase. Combining them covers each other's blind spot.

Contextual BM25 is the less obvious half. Adding a sentence of context to a chunk puts the document's key terms, the company name, the product, the year, into the chunk's own term statistics. That is why the lexical index gains from the same preprocessing as the dense one, and why the combination improves on either alone.

What this says about where production RAG landed

Read as a research result, this is a modest one: one company's internal benchmark, one embedding model family, one reranker, and no confidence intervals. Read as a description of practice it is more interesting, because every component is something teams had converged on separately. Hybrid dense plus lexical retrieval, a reranking stage that is cheap because it only sees a shortlist, generous top-k into the model, and some form of chunk enrichment. The note's contribution is to name the enrichment step, measure it, and show that it stacks with the others rather than replacing them.

The thing we would push on is the eval. A 5.7 percent failure rate at top 20 is already a strong baseline, and the queries that remain after each improvement are presumably the harder ones. Whether the same relative gains show up on a corpus of messy internal documents, where the chunker itself is the weak point, is a question the note does not answer and one we would like to see someone test on an open dataset.

What we would try next

Two experiments seem worth an afternoon. The first is to vary what the context generator sees, since giving Haiku the whole document may be unnecessary for long files, and a section heading plus the surrounding page might recover most of the gain at lower cost. The second is to check whether the context sentence is doing its work through the embedding or through BM25, by running contextual BM25 with plain embeddings and the reverse. If most of the gain comes through the lexical side, that changes which embedding model you should bother paying for.

Sources

  1. Anthropic, Introducing Contextual Retrieval