Two launches, one question

Within ten days of each other, two posts made opposite bets on how to retrieve documents. On July 21 Morphik published a piece titled Stop parsing docs, arguing that vision language models are now good enough to read a page image directly, with no OCR, layout detection or reconstruction. On July 30 Google announced general availability of gemini-embedding-001, a text embedding model positioned squarely for retrieval augmented generation and context engineering, with customer numbers to back it.

The question underneath both is what the unit of retrieval should be. For years it has been a chunk of text, produced by a pipeline that turns a PDF into strings. Morphik's claim is that the chunk is the problem. Google's launch is a reminder that the chunk is also where nearly all of the tooling, evaluation and money currently sits.

The case against the parse

Morphik's argument is about how many places a parsing pipeline can fail. They count seven steps, OCR, layout detection, reading order reconstruction, figure captioning, chunking, embedding and storage, and note that each one drops information the next cannot recover. The examples are ordinary. A thousand written as 1,000 comes out of OCR as l,0O0. The percentages from a chart legend end up scattered into unrelated paragraphs. A table loses its header row at a chunk boundary and every value beneath it loses its meaning.

The alternative is to embed the page image with a ColPali style model, in their case a SigLIP-So400m vision tower paired with PaliGemma-3B, and retrieve over the resulting multi-vector representation. Spatial relationships, table structure and chart context survive because nothing tried to flatten them. On their own 45 question evaluation over financial filings from NVIDIA, Palantir and JPMorgan, they report 95.56 percent accuracy against 72 percent for a LangChain pipeline with semantic chunking, about 67 percent for competing end to end providers, and 13.33 percent for OpenAI's file search. On ViDoRe they report 81.3 percent nDCG at 5 against 67.0 for traditional parsing.

The cost that vision retrieval has to pay

Multi-vector page embeddings are expensive to search. Morphik is candid that their first implementation took three to four seconds per query. What got them to 30 milliseconds was MUVERA, which collapses multi-vector similarity into a single vector comparison, and Turbopuffer as the store. So the honest version of the claim is that vision retrieval is fast now, after an engineering effort that most teams would not undertake, on top of a representation that is still larger per page than a text chunk.

The evaluation is also theirs. A 45 question set built by the vendor on documents they chose is a demonstration, and ViDoRe was designed by the ColPali authors around visually rich documents. We believe the direction of the result, since it matches what we have seen when tables and charts carry the answer. We would want a third party to run it on a corpus where most of the content is plain prose before we believed the size of the gap.

What the text side looks like at its best

Google's post is useful as a picture of what a well run text pipeline achieves. Box reports the correct answer over 81 percent of the time with a 3.6 percent lift in recall from switching to Gemini Embedding. Everlaw reports 87 percent accuracy over 1.4 million legal documents. Mindlid reports 82 percent top 3 recall at a 420 millisecond median, a 4 percent recall lift. re:cap reports F1 gains of 1.9 and 1.45 percent over previous Google models. Those are gains in the low single digits over an already strong baseline, which is what a mature technology looks like.

The model supports Matryoshka representation learning, so embeddings can be truncated with small quality loss, and it is multilingual. Nothing in the post mentions images. That is the gap Morphik is pointing at. The best text embedding model of the month is still downstream of a parser, and whatever the parser lost is gone before the embedding sees it.

Where we would draw the line

Our working rule after reading both is to split by where the answer lives. If the answer is in a sentence, text retrieval with a modern embedding model is cheaper, better tooled and easier to evaluate, and a few points of recall are not worth a new index type. If the answer is in a cell, a legend, a diagram or a position on a page, parsing is where the error enters, and page image retrieval is the only approach that does not have to reconstruct what the page already showed.

The experiment we would like someone to run is the failure overlap. Take a corpus with both kinds of content, run both pipelines, and report which questions each one gets wrong. If the sets are mostly disjoint, a router in front of both is the right product. If vision retrieval's failures are a subset of parsing's, the argument for parsing is over, and Morphik's title is correct.

Sources

  1. Morphik, Stop parsing docs (Jul 21, 2025)
  2. Google Developers Blog, Gemini Embedding: powering RAG and context engineering (Jul 30, 2025)