The claim that started it

On February 15, 2024, Google announced Gemini 1.5 Pro with a standard 128,000-token context window and an experimental preview of up to one million tokens, with research runs reaching ten million. In the needle-in-a-haystack test the model found the embedded text 99 percent of the time in blocks up to a million tokens. A million tokens is roughly an hour of video, eleven hours of audio, or more than 700,000 words in one prompt. Within days the reaction on social media, quoted by Zilliz, was Yao Fu writing that "the 10M context kills RAG".

We remember thinking at the time that the needle test was the wrong benchmark to draw that conclusion from. Finding one planted sentence in a haystack is a retrieval task the model was clearly good at. Reasoning across a haystack, or finding twelve needles that contradict each other, was a different question and nobody had a clean number for it. Two years later we want to write down what actually happened, because the debate did resolve, just not in the way either camp said.

What the retrieval side argued

The Zilliz response, which is representative of the vector database vendors, made four arguments it labelled velocity, value, volume and variety. On velocity it cited a 30-second response time for a 360,000-token context, which is fine for a batch job and unusable for a chat interface. On value it worked out that a million tokens at $0.0015 per thousand costs $1.50 per request. On volume it pointed out that ten million tokens is still tiny next to a search index. On variety it noted that time series, graphs and code diffs do not fit the "pour text into the window" model at all.

Those were correct in 2024 and mostly still are. Where the piece went wrong, in our reading, was the conclusion that vector retrieval specifically would persist the way hard drives persisted after RAM got cheap. Retrieval persisted. The vector part turned out to be optional.

What the agentic coding tools did instead

The clearest evidence came from a place nobody in the 2024 debate was looking at. Fabio Akita, writing this month, walks through what the leaked Claude Code source from March 31 shows about how the most heavily used coding agent handles memory and retrieval. There is no vector database in it. Memory is a markdown index of about 25 kilobytes plus topic files, and search over past transcripts is lexical grep over JSONL logs. A consolidation subagent runs grep to fold memories together. Retrieved memory is treated as a hint that has to be verified against the current state of the repository, never as ground truth.

That design is a direct answer to the failure modes of chunk-and-embed that Akita lists: false neighbours where cosine similarity returns something on the same topic that does not answer the question, chunk boundaries that split a table from its caption, index staleness whenever a file changes, and the opaque failure where you cannot tell why the wrong passage came back. A grep either matches or it does not, and it matches the file as it exists right now. Akita also cites BEIR results showing BM25 matching or beating dense retrievers out of domain, which lines up with what most teams found when they finally measured.

The models made this possible. Claude Opus 4.6 and Sonnet 4.6 both take a million tokens as of this spring, Gemini 3.1 Pro has experimental two-million modes, and the rest of the field sits at 200,000 to 400,000. When the window is that large, a lexical filter followed by loading whole files is a workable retrieval strategy, and the model does the fine-grained filtering that a reranker used to do.

How caching changed the arithmetic

The cost argument from 2024 depended on paying for the whole context on every call. Akita gives current numbers for a query with 200,000 input tokens and 2,000 output tokens: $0.63 on Sonnet 4.6, $3.15 on Opus 4.6, $0.42 on Gemini 3.1 Pro, $0.12 on GLM 5, and $3.36 on GPT 5.4 Pro. Those are still real money at volume. But with prompt caching, a cached follow-up on Sonnet drops to roughly ten cents, and he works out that 30,000 such queries cost around $3,200.

Set that against the cost of building and running a retrieval pipeline, which he estimates at 40 to 80 engineering hours before you count monitoring, re-embedding jobs and the on-call for when the index drifts. For a great many internal tools the tokens are cheaper than the engineers. That inverts the 2024 framing, where retrieval was the cheap option and long context was the luxury.

The practical recipe that falls out of this, and which Akita calls lazy retrieval, is to keep documents raw on disk, run a fast lexical filter such as ripgrep or BM25, load generously, let the model do the fine filtering, and only add embeddings if measured data on your corpus shows lexical search failing. We would put it slightly differently: start with grep and a big window, and treat a vector index as a performance optimisation you add once you can name the query class it fixes.

Where retrieval still wins outright

None of this means the window replaced the index. Akita lists the cases where classic retrieval is still the right call and they match our experience. Corpora in the hundreds of gigabytes, where no filter narrows things enough. Vocabulary that is scattered across synonyms, the customer-support case where users describe the same failure a hundred ways. Non-text modalities. Latency budgets under 100 milliseconds. And compliance settings where you need an audit trail of exactly which passages were shown to the model.

The research record over the two years is consistent with a split decision. Akita summarises an EMNLP 2024 paper from Google DeepMind finding long context beats retrieval on quality while retrieval saves tokens, with a self-routing hybrid as the practical answer, an ICML 2025 study concluding the winner depends on model, context size and task type, and Anthropic work from September 2024 showing that hybrid BM25 plus vector plus reranker beats pure vector. We have not re-run those evaluations ourselves and would want to before leaning on any single number from them.

What we would try next

The open question we care about is effective context. The advertised window is a million tokens. The window over which the model actually reasons well, on tasks harder than finding a planted sentence, is smaller and nobody publishes it. We would like to see a benchmark that plants twenty related facts across a document and scores the model on combining them, run at 50k, 200k, 500k and a million tokens, for each frontier model. Our expectation is a cliff well before the advertised limit, and that the location of the cliff, not the headline window, is what should drive the stuff-versus-retrieve decision.

If someone has that curve for the current models, we would trade a lot of the 2024 debate for it.

Sources

  1. Google: Our next-generation model, Gemini 1.5
  2. Zilliz: Will RAG be killed by long-context LLMs?
  3. Fabio Akita: RAG is dead, long live long context