A category gets priced in a week

Pinecone announced a $100 million Series B on April 26, led by Andreessen Horowitz, at a post-money valuation of $750 million. That brings the company to $138 million raised since 2021, on top of a $10 million seed and a $28 million Series A. TechCrunch reports around 100 employees today, a plan to reach 150 to 200 by the end of the year, and a customer count of about 1,500, up from a handful a year ago. Shopify, Gong, HubSpot and Zapier are named as customers.

The pitch is simple and the company states it plainly. Pinecone calls itself "Long Term Memory for AI". Peter Levine at a16z calls the vector database "a fundamental component in the new AI data stack". Edo Liberty, the CEO, says the adoption pattern looks like consumer growth applied to deep infrastructure, and that he has never seen anything like it. The announcement cites a $9 billion search infrastructure market and a $110 billion generative AI market as the prize.

Qdrant, Zilliz and Chroma are raising in the same window. So in the space of a few months a whole database category has been priced into the assumption that an LLM application needs a dedicated store for embeddings. We want to look at that assumption carefully, because we think it bundles three separate claims that do not stand or fall together.

What retrieval solved

The underlying need is real. A language model knows nothing about your documents, and the context window is too small to paste them all in. So you split the documents into chunks, embed each chunk, store the vectors, and at query time embed the question and pull back the nearest few chunks to put in the prompt. This is the pattern behind every "chat with your PDF" demo, and it works well enough that people ship it.

Pinecone and TechCrunch both list the advantages that follow from this arrangement. The data lives outside the model, so you can update it, delete it, and satisfy a GDPR request without retraining anything. The model sees only what it was given, which reduces the rate at which it makes things up. Semantic search finds a passage even when the wording differs from the query. None of that is controversial and none of it is where the disagreement lies.

Three claims hiding in one product

The first claim is that retrieval over your own data is necessary. We agree. The second claim is that the retrieval should be dense vector search with a cosine similarity top-k. That is one method among several, and it happens to be the one that a vector database is built to serve. The third claim is that this method needs a purpose-built, separately hosted database. That is the claim the $750 million is riding on, and it is the weakest of the three.

Here is the worked example that makes us nervous. Take a corpus of engineering documentation, chunk it into 500-token pieces, and ask "what is the timeout on the retry loop". The nearest chunks by cosine similarity are the ones that talk about timeouts and retries in general. The chunk that contains the actual constant may rank fifth or twelfth, because a single number does not move an embedding much. A keyword search for "retry" and "timeout" finds it immediately. Dense retrieval is good at paraphrase and bad at exact terms, and most real questions about private data contain exact terms.

Chunking has its own problems. The boundary you draw splits a table from its caption or a definition from its first use. The model then receives half of a thought and fills in the rest, which is the thing you added retrieval to prevent. None of this is fixed by a better index. It is a property of the chunk-and-embed recipe itself.

What we expect to survive

Our guess, and we want to be clear it is a guess in April 2023, is that the retrieval layer survives and the assumption that it must be a dedicated vector store does not. Postgres already has vector extensions. Elasticsearch and OpenSearch have had approximate nearest neighbour search for a while. For a corpus of a few hundred thousand chunks, a flat index in memory is fast enough that the database is not the bottleneck. The bottleneck is whether the right chunk was retrieved at all.

That points to where we think the engineering effort will go. Hybrid search that combines a lexical score with a dense score, so that exact terms and paraphrases both rank. A reranking stage that reads the top fifty candidates properly instead of trusting one dot product. Retrieval units that follow document structure rather than fixed token counts. And, for a lot of teams with a lot of code and text on disk, plain grep, because the corpus is small and the terms are known.

Pinecone may well be right that this is a $9 billion market. Managed infrastructure that scales to billions of vectors has a clear customer in large consumer search. What we doubt is that the median team building on top of a language model this year has a problem that a dedicated vector database is the right tool for. Their problem is retrieval quality, and the fix for that is mostly upstream of the index.

What we would want someone to measure

The experiment we would like to see published is boring and we have not seen it yet. Take three or four real internal corpora, write a hundred questions each with a known answer location, and compare recall at five for dense-only, lexical-only, and a simple hybrid. Then add a reranker to each and compare again. Our expectation is that hybrid plus rerank wins on most corpora, that lexical alone beats dense alone on technical text, and that the choice of vector store makes no measurable difference at that scale.

If someone has run that and the dedicated store wins, we want to read it. Until then the funding round tells us how much investors expect the category to be worth, and nothing about whether the method at the centre of it is the right one.

Sources

  1. TechCrunch: Pinecone drops $100M investment on $750M valuation
  2. Pinecone: Series B announcement