A Flash model that costs Pro money

Google announced Gemini 3.5 Flash at we/O on May 19, released without a preview tag, with a January 2025 knowledge cutoff, a 1,048,576 token input limit and 65,536 tokens of output. The price is $1.50 per million input tokens and $9 per million output tokens. Simon Willison worked through the comparisons on the day. That is three times the price of Gemini 3 Flash Preview and six times Gemini 3.1 Flash-Lite. Gemini 3.1 Pro sits at $2 and $12, which puts the new Flash within shouting distance of the model it is supposedly the cheap alternative to.

The per-token price is only half the story, and the half that matters less. Willison quotes Artificial Analysis's figures for the cost of running their full benchmark suite on each model. Gemini 3.5 Flash on its high reasoning setting cost $1,551.60. Gemini 3.1 Pro Preview cost $892.28. Gemini 3 Flash Preview with reasoning cost $278.26 and 3.1 Flash-Lite cost $93.60. On that measure the new Flash is roughly 5.6 times dearer than its predecessor and about 74 percent more expensive than Pro.

Why the benchmark bill diverges from the price list

The two comparisons disagree because they measure different things. The price list charges per token. The benchmark bill multiplies price by the number of tokens the model chose to emit, and a reasoning model on a high effort setting emits a great many. If 3.5 Flash writes several times as many thinking tokens per question as 3.1 Pro, it can be cheaper per token and more expensive per answer at the same time. That is what the Artificial Analysis numbers show.

This is why we have stopped reading price per million tokens as the cost of a model. The unit that matters to anyone running a workload is cost per completed task at a given accuracy, and that number now depends on a dial the provider sets and the caller can partly adjust. Two models with identical price lists can differ by a factor of five in what they cost to run on the same questions. Willison's own note that the 3.5 Flash benchmark run cost significantly more than 3.1 Pro is the cleanest illustration of that we have seen from a single vendor.

The trend this breaks

For about three years the reliable story was that the cost of a fixed level of capability fell steeply every year. Each generation of a small model matched the previous generation of a large one at a fraction of the price, and a lot of product planning quietly assumed that would continue. That story was about model size. Smaller weights, better data, better distillation, cheaper serving. It still holds along that axis. What has changed is that a second axis, tokens spent thinking per query, has opened up, and vendors are now spending along it.

Google is charging more for Flash because Flash now does more work per query, and because the product tier names no longer map to a fixed compute budget. The name says small and fast. The bill says it reasons at length. We would not be surprised if the Flash and Pro labels come to mean something closer to latency profile than cost, with the cost governed by effort settings that change from release to release.

What it does to the scaling-law story

The pretraining scaling laws related loss to parameters and training tokens, and the cost consequences were straightforward because inference cost scaled with parameters. Test-time compute adds a term that is spent at serving time and chosen per request. The plots that tie accuracy to thinking tokens look like scaling laws, and they behave like them in the sense that more compute buys more accuracy with diminishing returns. The difference is who pays and when. Training compute is paid once by the lab. Reasoning compute is paid on every call by the customer.

That shifts the economics in a way that the old curves did not capture. A lab can ship a model that is small by parameter count and expensive by reasoning budget, and the headline benchmark score will reflect the budget rather than the size. Comparing models on capability without holding cost per task fixed is now close to meaningless, and comparing them on price per token is worse than meaningless because it points the wrong way. The Artificial Analysis cost column is the one we now look at first.

What we would want measured

The experiment we want to see is a cost-matched comparison. Take 3.5 Flash at a lower effort setting that brings its benchmark bill down to 3.1 Pro levels, and see which one wins on accuracy at equal spend. If Flash still wins, the price increase is buying something and the trend of cheaper capability survives, just measured differently. If Pro wins at equal spend, the new Flash is a worse deal for most workloads and the name is doing a lot of work.

Either way, we think we should retire the assumption that next year is cheaper by default. Cost per task can go up with a release, and this month it did. Anyone budgeting for inference should be tracking tokens emitted per task on their own workload, per release, rather than reading the price list.

Sources

  1. Simon Willison: Gemini 3.5 Flash (May 19, 2026)