Predictable scaling: the one chart in the GPT-4 report that mattered
The GPT-4 technical report says nothing about architecture, data or compute, and then reports that final loss was predicted from runs using at most one ten-thousandth of the compute before the main run finished. That figure is the engineering claim of the release, and it says a lot about how frontier training is now planned.
What the report refuses to say
The GPT-4 technical report went up on March 15 and the sentence everyone quoted was the refusal. Citing both competition and the safety implications of large-scale models, it says the report contains no further details about the architecture, including model size, hardware, training compute, dataset construction, training method, or similar. For a document called a technical report that is a remarkable thing to write, and it has been read as the end of the era in which frontier labs published what they built.
We have been through the report several times since and we think the reaction to the silence has drowned out the one thing the report does disclose. Section 3 is titled Predictable Scaling. It is short, it has two figures, and it makes a claim about engineering practice that we would rank above any benchmark number in the document.
The claim
The section opens by explaining the motive. For a run the size of GPT-4 it is not feasible to do extensive model-specific tuning, so the team built infrastructure and optimisation methods with very predictable behaviour across scales. The test of that claim was a prediction. Before the main run had produced any partial results, they fit a scaling law of the form L(C) equals a times C to the power b plus c, with an irreducible loss term as in Henighan and colleagues, to models trained with the same methodology at up to 10,000 times less compute. They then used it to predict GPT-4's final loss on an internal codebase held out of training. Figure 1 shows the prediction landing on the final value.
The second figure does the same for a capability rather than a loss. Pass rate on a subset of HumanEval was predicted by extrapolating from models with at most 1,000 times less compute, using an approximate power law on the mean log pass rate. The report notes that individual HumanEval problems can get worse with scale, which is why they aggregate. And it shows one task from the Inverse Scaling Prize, hindsight neglect, where smaller models get worse with size and GPT-4 reverses the trend.
The abstract carries the 1,000 figure. The body carries the 10,000 figure for loss. Both are about extrapolating four orders of magnitude of compute from a family of small runs, and both were made before the answer was known.
Why this is the real result
Consider what has to be true for that prediction to work. The optimiser, the learning rate schedule, the data mix, the initialisation and the parallelism all have to behave the same way at every scale in the family, or the small runs tell you nothing about the big one. A team that can do this has turned a frontier training run from an experiment into a procurement decision. They know, before spending the compute, roughly what loss they will get and roughly what a coding benchmark will read. That is what makes it possible to commit the sort of budget these runs require.
It also changes what the small runs are for. In the older practice, small models were where you tried ideas and the big model was where you hoped they transferred. Under predictable scaling, small models are instruments for measuring the big one, and an idea that does not fit the curve is a problem to fix in the infrastructure rather than a finding. The report says as much when it lists building a stack that scales predictably as a large focus of the project.
The report's own gloss is about safety. It says that accurately predicting future capabilities is important for safety, and that the team plans to register performance predictions across capabilities before large training runs begin, with the hope that this becomes a common goal in the field. We read that as a commitment worth holding them to. A registered prediction, published before the run, is exactly the kind of evidence that would let an outside party check whether a lab understands what it is building.
What is missing
The section proves less than it might seem to. The loss prediction is on an internal codebase and the capability prediction is on a subset of HumanEval. There is no statement of how many small runs went into each fit, how far the prediction was from the final value in absolute terms, or which capabilities they tried to predict and could not. The hindsight neglect figure is a reminder that some behaviours do not follow the curve, and the report does not say how the team decides which ones will.
There is also a tension between the two halves of the report. Predictable scaling is a method that anyone with the same infrastructure discipline could reproduce, and it would be the most useful thing in the document to publish in detail. It is precisely the kind of detail that the refusal paragraph rules out. What we are given is the shape of the curve and the assertion that it worked.
What we would want from every lab
The practice we want to see adopted is the one the report proposes and does not fully perform. Before a large run, publish the scaling fit, the metrics it predicts and the predicted values, in a form that is time-stamped and cannot be edited. After the run, publish the actuals. Any lab that can do what Section 3 describes can do this at no extra cost, and it would tell the rest of us far more about the state of the field than a table of exam scores.
For a small lab the implication is different. If frontier runs are planned from families of small runs, then the science of those families, of what transfers and what does not across four orders of magnitude, is something we can work on with the compute we have. The GPT-4 report shows that the curve exists. It does not show why, and that is a question open to anyone.
Sources
From the foundation