A paper written as a rebuttal

Jimmy Lin's group at Waterloo, with Tommaso Teofili at Roma Tre, posted a short paper last week whose title does most of the work. They took OpenAI's ada2 embeddings for the MS MARCO passage collection, indexed them with the HNSW support that already ships in Lucene, and ran the standard evaluation. The retrieval quality was fine. The stated goal is to challenge what they call the prevailing narrative, which is that dense retrieval requires a dedicated vector database as a new component in the stack.

We read it as a sober note pinned to the door of every startup pitch we have seen this year. The paper names Pinecone and Weaviate as examples of the category and quotes the broader claim that vector databases are now must-have infrastructure. Its counterclaim is narrow and worth stating precisely. It does not say Lucene is faster or better. It says that since search is a brownfield application where organisations already run Elasticsearch, OpenSearch or Solr on top of Lucene, adding a second system has a cost that the benefits do not obviously cover.

What they actually built

The setup is deliberately ordinary. Passages and queries were embedded with ada2, which accepts up to 8191 tokens of input and produces 1536-dimensional vectors. The vectors were indexed in Lucene 9.5.0 through Anserini, the group's own toolkit, and evaluated on the MS MARCO development queries plus the TREC 2019 and 2020 Deep Learning Track queries. The embeddings are downloadable, so anyone can reproduce the numbers without paying OpenAI for the passage encodings.

On the dev set, ada2 reaches an RR@10 of 0.343 and R@1k of 0.984. For comparison, the paper's own table lists BM25 at 0.184 and 0.853, TAS-B at 0.340 and 0.975, and ColBERT-v2 at 0.397 and 0.984. On DL19 the nDCG@10 is 0.704 and on DL20 it is 0.676, which puts ada2 roughly level with TCT-ColBERTv2 and a little below SPLADE++ ED. The authors' reading is that a commercial embedding endpoint gives you a strong, if not state of the art, bi-encoder with no training on your side.

One detail that made us smile. Lucene 9.5.0 could not index the vectors as shipped, because its HNSW implementation limited vectors to 1024 dimensions. The authors had to work around it and note that at the time of writing there is no public Lucene release that indexes ada2 vectors directly. That is the kind of friction the vector database vendors are selling their way past, and the paper is honest that it exists.

The performance section is the honest part

The reported throughput is what we would want anyone quoting the title to also quote. Indexing the passage collection took around three hours with 16 threads, with M set to 16 and efConstruction set to 100. Query evaluation with 16 threads and efSearch at 1000, retrieving 1000 hits per query, ran at 9.8 queries per second on a dated two-socket Xeon server with 1 TB of RAM. The vectors themselves were stored as gzipped JSON text at 109 GB, against a theoretical 54 GB for raw floats, which the authors describe as inefficient but convenient.

The discussion section goes further and cites a separate benchmark by Ma and colleagues that found Lucene 9.5.0 achieves around half the query throughput of Faiss under comparable settings, while scaling better across threads. The authors' phrase is that Lucene is unequivocally slower. Their argument for staying with it is that Faiss is mature and has limited headroom, whereas Lucene has many places left to improve and, in their view, strong signs of commitment from Elastic and Amazon to do that work.

So the claim is a cost-benefit judgement rather than an engineering win. If you run a search platform already, HNSW in Lucene is adequate and the operational cost of a second store is real. If you are starting from nothing, or if you need tens of thousands of queries per second against a billion vectors, nothing in this paper argues that Lucene is the right tool today.

Where we land

We agree with the framing more than we expected to. Most retrieval-augmented systems we have looked at this year have corpora in the hundreds of thousands to low millions of chunks, query rates that a single machine handles, and a team that already has to keep a keyword index running for filtering and exact match. Adding a vector database to that stack introduces a second source of truth, a second set of failure modes, and a second bill. The paper's point that relational databases stayed a fixture through every wave of alternatives is a fair analogy.

The paper also gestures at Vespa as an example of a platform that handles both dense and sparse retrieval natively, and at hybrid setups built inside Elasticsearch. Those seem like the practical middle ground for the next year or two. The bet we would not make is that the dedicated stores disappear. Their advantage is at the scale and latency Lucene cannot yet reach, and the paper concedes that gap exists.

What we would like someone to run is the experiment this paper stops short of. Take the same MS MARCO embeddings, index them in Lucene, Faiss and two commercial vector stores, and report quality, throughput and the number of engineer hours to get each one serving with metadata filtering. The quality column will be flat. The other two columns are where the actual decision lives, and nobody with a product to sell has published them.

Sources

  1. Lin, Pradeep, Teofili, Xian: Vector Search with OpenAI Embeddings: Lucene Is All You Need (arXiv 2308.14963)