The claim

Matei Zaharia, Omar Khattab and nine coauthors published an essay on the BAIR blog on February 18 with a simple thesis. State-of-the-art results in AI are increasingly produced by compound systems, which they define as systems that tackle a task using multiple interacting components, including multiple calls to models, retrievers or external tools. The contrast is with a single statistical model, trained end to end, that you query once and trust.

The examples they list are the ones we would have chosen too. AlphaCode 2 generates up to a million candidate programs, runs them, clusters them and scores them, and lands at the 85th percentile of human competitors. AlphaGeometry alternates between a language model proposing constructions and a symbolic engine checking deductions. Medprompt wraps GPT-4 in nearest-neighbour example retrieval, model-generated chains of thought and an ensemble of up to eleven samples, and beats Med-PaLM with a general model. Gemini's headline 90.04 percent on MMLU comes from a CoT@32 procedure that samples 32 reasoning chains and returns the majority answer only when enough of them agree, against 86.4 percent for GPT-4 five-shot.

Why the authors think this happened

The essay gives four reasons and the first is the one that matters for anyone with a fixed compute budget. Some tasks are easier to improve with system design than with scale. Their AlphaCode-style example is that sampling many candidates and testing them can move a task from 30 percent to 80 percent, while a bigger model might buy you five points. Iterating on a pipeline takes hours. Waiting for a training run takes months.

The other three are about things a frozen model cannot do at all. Training data has a cutoff, so any application that needs today's facts, or per-user access control on what the model may see, needs a retrieval component. Neural networks do not come with behavioural guarantees, so filtering outputs and attaching citations is easier to do outside the model than inside it. And each model sits at one point on the quality-cost curve, while a system can route cheap requests to a small model and expensive ones to a big one. The essay cites Databricks data that 60 percent of LLM applications use some form of retrieval augmentation and 30 percent use multi-step chains.

What this does to evaluation

Here is the part that bothers us. When Gemini reports 90.04 percent on MMLU, that number belongs to a procedure, and the procedure includes a sampling budget, an agreement threshold and a fallback rule. The same report shows Gemini at 83.7 percent five-shot. Both numbers are real. Neither one is the model's score in the way we used to mean it. Any comparison across labs now has to state the inference procedure or it is not a comparison.

The same problem shows up in every retrieval system. If a RAG pipeline gets a question right, was that the retriever or the generator? If it gets one wrong, which component do you retrain? The essay is honest that the design space is enormous even for the simplest pipeline, and gives a nice example of the budgeting problem. If you have 100 milliseconds to answer, do you spend 20 on retrieval and 80 on generation, or the reverse? Nobody has a principled answer yet, and the benchmarks we publish do not record the choice.

The essay's proposed fix is to treat the pipeline as the object of optimisation. DSPy, which Khattab leads, tunes prompt instructions, few-shot examples and other per-module parameters against a metric, the way PyTorch tunes weights. FrugalGPT learns a routing policy and is reported to match the best single service at up to 90 percent lower cost. We like both, and we would add that they only work if the metric is good, which pushes the hard problem into evaluation rather than removing it.

Who owns the result

There is a quieter shift under the technical one. When the artefact was a model, it had one owner, one training run and one card. A compound system might use a hosted model from one company, an embedding model from another, a vector store, a code sandbox and a scoring model, with prompts and glue written by the team deploying it. The essay lists the missing operations tooling, logging of long traces, data pipelines feeding the vector store, and security risks it calls unforeseen. All of that is a way of saying that responsibility for behaviour has spread across parties that do not share a bug tracker.

For research this cuts both ways. A small group cannot train a frontier model, but it can build a compound system around one and beat the model on a narrow task, which is roughly what Medprompt showed. That is good news for labs like ours. The bad news is that a result built this way is harder to reproduce, because it depends on an API version, a retrieval corpus snapshot and a sampling budget that may not be published, and the underlying model can change under you.

What we would want to see next

The obvious next step is to report inference procedure alongside score as a matter of course, the way we report training compute. Samples drawn, agreement rules, retrieval corpus, tool calls permitted. If the number is a property of a system, the system should be described well enough to rebuild.

We also want ablations that switch components off one at a time, published with the headline result. The compound framing is useful precisely because it lets us ask which part of the pipeline did the work. If we are going to accept 90 percent from a procedure rather than from a model, we should at least learn where the extra points came from.

Sources

  1. BAIR blog: The Shift from Models to Compound AI Systems (February 18, 2024)