The GPT-4 technical report and the paper that told us nothing
OpenAI released a hundred-page technical report for GPT-4 that states, in one paragraph, that it will say nothing about architecture, size, hardware, compute, data or training method. Reading notes on what is in the document, what is missing, and what an independent lab can still take from it.
The paragraph
Section 2 of the report is titled Scope and Limitations of this Technical Report, and it contains the sentence everyone has been quoting. Given both competitive considerations and the safety implications of large-scale models like GPT-4, the report contains no further details about the architecture, including model size, hardware, training compute, dataset construction, training method, or similar. The preceding sentences tell us only that GPT-4 is a Transformer-style model, pretrained to predict the next token on public and licensed data, and fine-tuned with RLHF.
The report then says OpenAI is committed to independent auditing and plans to make further technical details available to additional third parties who can advise on how to weigh competition and safety against the scientific value of transparency. The authorship line asks that the work be cited as OpenAI (2023), with individual contributions listed in an appendix. The document runs to a hundred pages. It has a bibliography, figures, tables of benchmark results, and no method.
What is actually in it
Strip out the withheld parts and three things remain. First, capability numbers. GPT-4 scores 86.4 percent on MMLU against 70 percent for GPT-3.5 and 75.2 percent for the best prior published result. It gets 67 percent on HumanEval against 48.1 percent. On a simulated Uniform Bar Examination it lands in the top 10 percent of test takers. The report also includes a contamination analysis for these benchmarks, estimating for example that around 25 percent of HumanEval overlaps with training data and reporting a small adjusted difference.
Second, predictable scaling. The report says a large focus of the project was a training stack whose behaviour could be forecast from small runs. Final loss on an internal codebase was predicted from models using at most ten thousand times less compute, and the abstract says some aspects of performance were predicted from models trained with no more than a thousandth of the compute. They registered a prediction for HumanEval pass rate before training finished, using a power law fit on a subset of problems, and it held. They also show the Hindsight Neglect task from the Inverse Scaling Prize, where performance had fallen with scale across earlier models and GPT-4 reverses the trend.
Third, safety. The report states that OpenAI spent six months on safety research, risk assessment and iteration before launch, describes adversarial testing with domain experts and a model-assisted safety pipeline, and notes that forecasters they consulted predicted that delaying deployment by a further six months would reduce acceleration risk. The system card that accompanies the report carries most of this material.
The moment papers stopped being papers
A paper, in the sense the field has used for decades, is a document that lets a competent reader reproduce or at least check the result. This document does not attempt that. It reports numbers from an artefact whose construction is secret and asks the reader to trust the numbers. The GPT-3 paper in 2020 gave the architecture, the parameter count, the data mixture and the compute. The PaLM paper last year gave all of that plus training details. The gap between those documents and this one is the gap between a scientific report and a product announcement with an appendix of evaluations.
We do not think the authors would disagree. The report is explicit that the omission is a choice, and it gives the reasons. Competition and safety. Both are real. Neither changes what the document is. A hundred pages of results with the method removed is a benchmark table, and a benchmark table produced by the people who built the system, on benchmarks they chose, with contamination they estimated themselves, is evidence of a particular and limited kind.
What worries us more than this report is the precedent. If the strongest model in the world can be announced this way and the field treats the document as a paper, then the next lab has no reason to do otherwise. The disclosure norm was already eroding. This is the moment it became optional at the frontier.
What an independent lab can still use
The predictable scaling section is worth more than the rest of the report combined, and it is worth something even without the details. It tells us that at least one lab can now forecast final loss from runs four orders of magnitude smaller, and can forecast a downstream metric well enough to bet on it in advance. That is a claim about the maturity of the engineering, and it is checkable in principle by anyone with a small-scale suite. A lab that cannot do this on its own runs is behind, and the report tells you what the target is.
The Hindsight Neglect result is a small but genuinely informative data point. A task that got worse with scale across every model anyone had tested got better at the frontier. Whether that is scale, RLHF, or data is exactly the kind of question the report cannot help with, but the observation itself constrains theories of inverse scaling, and it is one we could not have made ourselves.
The contamination analysis is a model of what to ask for. The authors estimated overlap between each benchmark and their training data, and reported the score on the uncontaminated portion alongside the raw score. That is more than most academic papers do. It is also a reminder that only the lab with the data can do it, which is one more reason the data should be described.
What we would do about it
The response that makes sense to us is to stop treating documents like this as papers and start treating them as claims to be tested. The benchmarks in the report are public. The model is behind an API. A third party can rerun MMLU, HumanEval and the bar exam under its own protocol and publish the differences, and the more of that happens the less the self-reported numbers matter. Where there are no numbers, on data and compute, the honest position is that we do not know, and any downstream analysis that assumes a parameter count is guessing.
The other response is to keep publishing the way we always have. Every open training run with a full method section is a data point that the field can check, and it is a standing argument that the alternative was a choice.
Sources
From the foundation