The pricing change

Anthropic shipped prompt caching on its API this week. The mechanism is simple. You mark a prefix of the prompt, typically the system prompt, tool definitions, and any large reference material, and the first request that sends it pays a cache write price. Subsequent requests that begin with the identical prefix pay a cache read price for those tokens instead of the normal input price. Writing costs 25 percent more than a standard input token. Reading costs 10 percent of it.

In dollar terms, on Claude 3.5 Sonnet at 3 dollars per million input tokens, a cache write is 3.75 dollars per million and a cache read is 30 cents. On Claude 3 Opus the same figures are 18.75 dollars and 1.50 dollars against a 15 dollar base. On Claude 3 Haiku they are 30 cents and 3 cents against 25 cents. The output price does not change. The break-even is at two requests. If a prefix is sent more than once, caching is cheaper.

What it makes affordable

The announcement carries three worked examples and they are worth reading as a guide to which designs just became viable. Chatting with a book loaded into context, about 100,000 tokens, drops from 11.5 seconds of latency to 2.4 seconds and costs 90 percent less per turn. A many-shot prompt of 10,000 tokens goes from 1.6 seconds to 1.1 and costs 86 percent less. A multi-turn conversation that keeps re-sending its history goes from around 10 seconds to around 2.5 and costs 53 percent less.

The book example is the one that changes behaviour. Before this, stuffing a 100,000 token document into the prompt for every question was something you did in a demo and then replaced with retrieval in production, because paying for 100,000 input tokens per query was not sustainable. At 30 cents per million on the cached read, 100,000 tokens costs three cents per query. That is within the range where you can simply hand the model the whole document and stop worrying about whether your retrieval pipeline found the right chunk.

Many-shot prompting is the second design that moves from expensive to routine. The finding over the past year has been that a few hundred in-context examples often beat a handful, but a few hundred examples is tens of thousands of tokens on every call. If those examples sit in the cached prefix, the marginal cost of including them drops by an order of magnitude and the latency penalty mostly disappears, because the cached prefix is not reprocessed.

The third design is the agent with a large fixed toolkit. Tool definitions, long instruction sets, and the corpus an agent searches over are all static across a session. An agent making twenty tool calls in a loop previously paid full price for its system prompt twenty times. Now it pays once and reads nineteen times at a tenth.

What it rewards

Caching is prefix-based, which means the design habit it rewards is ordering. Everything that does not change between requests goes first. Everything that does change goes last. A system prompt that interpolates the current date, the user's name, or a session identifier at the top breaks the prefix on every call and gets no cache hits. The same prompt with those values moved to the end of the user turn gets nearly full hits.

It also rewards stability over time. A team that edits its system prompt daily will pay the write premium daily, which is fine, but a team that generates the system prompt dynamically per request, with reordered sections or per-user variations early in the prefix, will pay the write premium on every single call and end up 25 percent worse off than before. The announcement frames caching as suited to workloads where a large amount of context is sent once and referred to repeatedly, and that phrase should be read as a requirement, not a suggestion.

The other thing it rewards is putting more in. The old discipline of trimming the system prompt to the bone because every token cost money on every call is now partly obsolete for the static portion. Instructions that were cut to save cost can go back in. Full documentation for a tool can replace a summary. The trade is that longer prefixes are more expensive to write on a miss, so the calculus is the hit rate, and a prefix that is hit hundreds of times can be as long as the context window allows.

What it punishes

The obvious loser is a request pattern with low reuse. One-off calls, batch jobs over many unrelated documents, and anything where the large part of the prompt is the varying part get no benefit and pay the premium if you turn caching on without thinking. The announcement's own multi-turn example is instructive here. It saves 53 percent rather than 90, because the history grows each turn and the newest turns are always uncached.

The subtler loser is retrieval-augmented design that puts the retrieved passages before the question. Retrieved content varies per query, so if it sits early in the prompt it invalidates the prefix. Retrieved passages have to move after the static block, which is a change to prompt templates that many teams will need to make deliberately.

Where we think this goes

The practical experiment we want to run is the one the book example implies. Take a question answering task over a document set where we currently use retrieval, load the whole corpus into a cached prefix where it fits, and compare accuracy and cost per query against the retrieval pipeline. Our expectation is that for corpora under the context limit the cached full-context version wins on accuracy and is now competitive on cost, and that the case for retrieval narrows to corpora that do not fit.

The pricing also creates an incentive we expect other providers to copy, because it makes the long context window a feature that people actually use at volume rather than a number in a launch post. Once cache reads are a tenth of input price everywhere, the interesting question is how long a cached prefix can live and what it costs to keep it warm, and that is where the next round of differences between APIs will show up.

Sources

  1. Anthropic, Prompt caching with Claude (August 2024)