GPT-5 and the expectations gap
GPT-5 was sold on August 7 as a significant step toward AGI and received as a refined product. Users mourned GPT-4o's tone and OpenAI brought it back within days. We think the disappointment was a measurement problem: nobody in the field can say any more what a big jump would look like.
What was promised and what arrived
The launch livestream on August 7 came with the largest framing OpenAI has used for a model. Sam Altman called GPT-5 a significant step along the path to AGI, described it as having PhD-level ability across a wide range of tasks, compared the effort to the Manhattan Project, and said the model made him feel useless. The day before, he posted an image of the Death Star with no caption. Whatever the model turned out to be, it was going to be judged against that.
What arrived was a system with a fast model, a deeper reasoning model, and a router that picks between them in real time, plus claims of state-of-the-art results on maths, coding, finance and multimodal benchmarks, faster responses, better health answers and lower hallucination rates. Early testers reported strong coding and maths and gains that felt incremental over GPT-4. Grace Huckins at MIT Technology Review called it above all else a refined product, and said it fell far short of the transformative framing. John Herrman at New York magazine wrote that casual users were unlikely to notice much difference, while developers and corporate users would.
The week the users pushed back
The reaction that dominated the first week had little to do with benchmarks. OpenAI removed the older models, including GPT-4o, from the picker for non-Pro users at launch, without warning. The router misbehaved on launch day, so many people got the fast model when they expected the reasoning one and concluded that GPT-5 was worse than what it replaced. Altman later said the model would seem smarter starting today once the autoswitcher was fixed, which is a sentence we do not think any previous launch has needed.
Then came the tone complaints. Users described GPT-5 as flat, uncreative and lobotomised, and one line that went around compared it to an overworked secretary. Kyle Orland at Ars Technica found GPT-4o a little more detailed and a little more personable. Altman acknowledged that OpenAI had underestimated how much people valued 4o's warmth, restored 4o as an option for Plus subscribers, said on August 13 that personality changes were coming, and shipped an update on August 15 meant to make the model feel warmer. A month of product work, compressed into eight days, driven entirely by how the model sounds.
Why the gap was a measurement problem
Here is the reading we find most convincing. The disappointment tells us less about the model than about the field, which has lost its shared unit for what a big jump is. In 2023, GPT-4 could be compared to GPT-3.5 on a bar exam and a set of academic benchmarks, and the jumps were large enough that nobody needed to argue about the metric. Two years later, the benchmarks that GPT-4 made famous are saturated, the ones that replaced them are contested, and the ones people actually feel, like tone and whether a coding agent finishes the job, have no agreed scale at all.
So a launch has to make its case with three incompatible kinds of evidence. Benchmark deltas, which are real but small in absolute terms and which most users never see. Vibes, which are immediate but which move in the opposite direction when the model becomes more careful. And a story about AGI, which is untestable by construction. GPT-5 did well on the first, badly on the second, and was sold on the third. The expectations gap is what you get when the marketing lives on an axis the measurement cannot reach.
The 4o episode is the cleanest evidence for this. A meaningful fraction of the paying user base preferred an older model that scores worse on every published benchmark. Either those users are wrong about what they want, which seems unlikely given that they are paying, or the benchmarks are not measuring what those users are buying. If a lab cannot tell from its evaluations that a large group of customers will revolt at a personality change, the evaluations are missing an axis that matters commercially, and probably one that matters for safety too.
Where the measurement could go
There are practical things to try. Preference evaluations already exist, and this launch is an argument for running them at scale before removing a model rather than after. A lab could publish, alongside the benchmark table, a held-out preference rate against the model being retired, broken down by task type. A warmth number would be a strange thing to see on a model card, but it would have predicted this week better than the hallucination rate did.
The security results from the first day also belong in this conversation. Two red-teaming firms, NeuralTrust and SPLX, reported getting the model to produce detailed instructions for making explosives within a day of release, and judged it unsafe for corporate use without further work. Those findings did not get much attention next to the tone complaints, which is itself a symptom. When the discourse cannot agree on how to measure capability, the loudest measurable thing wins, and this week the loudest measurable thing was a personality.
We would want someone to try a simple exercise before the next frontier launch: write down, in advance and in public, what result would count as a big jump and what would count as incremental. If the labs cannot do that, the expectations gap is guaranteed no matter what the model does. If they can, we might finally learn something from the reaction.
Sources
From the foundation