The claim

Omar Khattab and twelve co-authors at Stanford and elsewhere posted DSPy to arXiv on October 5. The claim is that hand-written prompt templates are the wrong abstraction for building systems out of language models, in the same way that hand-written assembly is the wrong abstraction for building software. The paper proposes a programming model where a pipeline is a graph of text transformations, each step is a module with a declared signature, and the actual prompts are produced by an optimiser rather than typed by a person.

We had spent the previous year maintaining a set of prompts for a retrieval pipeline that broke every time we changed the model. The idea that the prompt could be a compiled artifact, regenerated when the model changed, was the first proposal we had seen that treated our problem as an engineering problem rather than a craft.

Signatures, modules, teleprompters

A signature is a natural language declaration of what a step does, written as input fields mapping to output fields. The example in the paper is question -> answer. It says what the transformation is and nothing about how the model should be asked to do it. A module implements a signature. Predict is the plain version. ChainOfThought inserts a reasoning field before the output. ReAct wraps a tool-using loop. Retrieve pulls passages. Modules are parameterised, and the parameters include the demonstrations that end up in the prompt and, optionally, the weights of a small model.

The optimisers, which the paper calls teleprompters, are where the compiling happens. BootstrapFewShot runs the pipeline on training inputs, keeps the traces that produced correct outputs according to a metric, and uses those traces as demonstrations in the prompt. That is rejection sampling turned into prompt engineering. BootstrapFewShotWithRandomSearch adds a search over which demonstrations to keep. BootstrapFinetune uses the bootstrapped traces to finetune a smaller model instead of prompting a larger one. The paper describes compilation as three stages, candidate generation by running the program, parameter optimisation, and higher-order optimisation such as building ensembles, and says it takes minutes to tens of minutes.

The numbers

On GSM8K, GPT-3.5 with a bare Predict module scored 25.2 percent. Compiling a ChainOfThought program with bootstrapped demonstrations raised that to 72.9, and an ensemble reached 81.6. Llama2-13b-chat went from 9.4 percent to 43.7 with a bootstrapped ensemble. On HotPotQA, a compiled multi-hop program scored 48.7 with GPT-3.5 and 42.0 with Llama2-13b-chat, and in both cases the ensemble was slightly worse than the single compiled program. The abstract summarises this as gains of more than 25 percent for GPT-3.5 and more than 65 percent for Llama2-13b-chat over standard few-shot prompting, and gains of 5 to 46 and 16 to 40 percent over pipelines using expert-written demonstrations.

Two things in the table matter more than the headline gains. First, the bootstrapped demonstrations beat human-written ones, which is the result that justifies the whole approach. If a person could write better prompts than the optimiser, compilation would be a convenience. Since they could not, it is a capability. Second, a 770M T5 and a 13B Llama became competitive with prompted GPT-3.5 on these tasks after compilation, which is the argument that the method changes which model you need rather than only how well one model does.

What compilation actually buys

The practical property is that a program can be recompiled. When the model changes, you rerun the optimiser against the same signatures and metric and get new prompts suited to the new model. The pipeline logic is untouched. That is the promise, and in our own use it has mostly held for the demonstrations part. The failure mode is the metric. Bootstrapping keeps traces that score well, so if the metric can be satisfied by the wrong reasoning, the compiler will happily fill the prompt with examples of it. Writing a good metric turned out to be as much work as writing a good prompt, with the difference that the metric survives a model change.

The other thing the paper gets right is that the prompt stops being the place where knowledge about the task lives. It lives in the signature, the metric and the training examples, all of which are inspectable. A hand-tuned prompt encodes the same knowledge as a pile of phrasing tricks that nobody can audit.

How the idea aged

Reading the paper again from a later vantage, the part that aged worst is the ensembling. Running five compiled programs and voting was a reasonable way to squeeze GPT-3.5 in 2023 and an expensive one to squeeze a model that can reason on its own. The part that aged best is the framing. Each new model generation broke hand-written prompts, and each time the teams we know with signatures and metrics recompiled in an afternoon while the teams with prompt files spent a week. The specific optimisers changed. The habit of separating what the step does from how the model is asked did not.

The open question is whether the abstraction survives models that need almost no prompting at all. If a signature and a single instruction get you within a point of the compiled version, the optimiser has nothing to optimise, and DSPy becomes a typed interface with a metric attached. That is still a good way to build a pipeline. We would like to see the GSM8K and HotPotQA experiments rerun on current models with a report of how much of the gap the compiler still closes, because that number is the honest measure of whether prompts still need compiling or whether the models have quietly compiled them for us.

Sources

  1. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines (arXiv 2310.03714)
  2. DSPy paper, full text (arXiv HTML)