GPT-5.4 and the one million token default
OpenAI shipped GPT-5.4 on March 5 with a 1,050,000 token context window, native computer use and tool search, then mini and nano variants twelve days later. An explainer on what has to change in attention, positional encoding and serving before a million tokens stops being a stunt.
What shipped
GPT-5.4 arrived on March 5 as the first mainline OpenAI model with native computer use, a tool search feature for large tool sets in the API, and a context window the model page lists at 1,050,000 tokens with 128,000 tokens of output. The mini and nano variants followed on March 17. The base model costs 2.50 dollars per million input tokens and 15 dollars per million output, mini costs 0.75 and 4.50, and nano 0.20 and 1.25. Simon Willison worked out that describing every one of his 76,000 photos with nano would cost about 52 dollars.
The detail that tells you the million tokens is still expensive to serve is in the pricing footnote. Requests that go beyond 272,000 tokens are billed at twice the input rate and one and a half times the output rate for the whole session. So the long window exists, and it is priced as something the provider would rather you used deliberately. That is fair, and the rest of this post is about why.
Attention is the easy part to explain
Standard attention compares every query position with every key position, so the work grows with the square of the sequence length. Doubling from 500,000 to a million tokens quadruples the attention cost of a full prefill. Modern kernels compute this in tiles and never write the full score matrix to memory, which removes the quadratic memory term, but the quadratic compute is still there. The practical answers are all forms of not attending to everything. Some layers attend only within a local window, some attend sparsely to a selected subset, and some architectures replace a fraction of attention layers with recurrent or state-space blocks that carry a fixed-size summary forward.
None of the public material on GPT-5.4 says which of these OpenAI uses, and we are not going to guess. What we can say is that a model advertising a million tokens at a price that is only twice the short-context input rate is telling you that its per-token cost at long range is much lower than dense attention would imply. Either the architecture is sparse in some dimension, or the serving stack is amortizing the cost across a large fleet in ways a single-request view cannot see.
Positional encoding decides whether the window is real
A model can accept a million tokens and still be useless past the length it was trained on. Rotary position embeddings encode relative distance by rotating query and key vectors at a set of frequencies. Positions beyond the training range produce rotations the model has never seen, and quality falls off. The standard fixes rescale the rotation frequencies so that a longer sequence maps onto the range the model knows, then continue training at longer lengths so the model learns to use the extra room.
This is the part that separates a context window from a marketing number. A window of 1,050,000 tokens only means something if the model retrieves and reasons over content near the start of that window as well as it does over content near the end. The published material we have read does not include long-context retrieval or reasoning evaluations for GPT-5.4 at the full window, and until someone publishes them we would treat the million as an upper bound on what you can send rather than a guarantee of what the model can use.
Serving is where the cost lives
Every token in the context leaves a key and a value vector in every attention layer, and that cache has to sit in accelerator memory for the duration of the request. At a million tokens the cache for a large model is measured in gigabytes per request, and it competes with the weights for the same memory. The pricing structure follows directly. Cached input at a tenth of the fresh input price rewards you for keeping a stable prefix so the provider can reuse the cache across calls. The long-context surcharge above 272,000 tokens reflects that a request holding that much cache blocks memory that could serve several smaller requests.
Tool search fits the same picture. If an agent has hundreds of tools, putting every definition into the prompt on every turn spends context on descriptions the model will not use. Letting the model search for the tool it needs keeps the resident context small, which keeps the cache small, which keeps the request cheap. Computer use pushes in the other direction, since screenshots and action histories accumulate quickly, and we expect the long-window pricing to matter most for exactly those agentic sessions.
What ordinary would look like
Ordinary means the long window stops being a separate product with a separate price. That requires three things to keep improving at once. Attention cost per token at long range has to keep falling, through sparsity or hybrid architectures, so that the compute premium disappears. Positional handling has to be validated with evaluations that test the full window, so that users can trust content anywhere in the prompt. And the memory footprint of the cache has to drop, through quantization of the cache or architectural sharing of keys and values across heads, so that a million-token request no longer displaces a dozen short ones.
The mini and nano releases are the clearer signal about direction. Simon reports that nano at maximum reasoning effort beats the previous GPT-5 mini, and that the new mini runs about twice as fast as the old one. Cheap, fast models with a large window are what make long-context use routine, because the cost of stuffing a repository or a day of logs into the prompt stops mattering. The experiment we want to see run on GPT-5.4 is a needle and reasoning sweep at 100,000, 500,000 and a million tokens, at all three model sizes, with the surcharge included in the cost column.
Sources
From the foundation