Two models and a switch

GPT-5 launched on August 7 to every ChatGPT user, with higher limits for Plus and unlimited use for Pro. The system card describes it as a unified system made of a fast, high-throughput model, a deeper reasoning model, and a real-time router that decides which one to use based on conversation type, complexity, tool needs and explicit intent. Underneath are gpt-5-main and gpt-5-main-mini on the fast side, and gpt-5-thinking, gpt-5-thinking-mini and gpt-5-thinking-nano on the reasoning side. API users bypass the router and set reasoning effort directly, with a new minimal setting that turns off most reasoning so tokens start streaming immediately. Pricing is 1.25 dollars per million input tokens and 10 dollars per million output for the full model.

The router is the interesting part. Since o1 last September, test-time compute has been a knob the user turns, by picking a model or a reasoning level. GPT-5 moved the knob inside the product. The model, or a component trained alongside it, looks at your message and budgets its own effort. If that works, most users never think about it again, and OpenAI saves an enormous amount of compute on the questions that never needed a reasoning model.

Launch day

It did not work on day one. The next day Sam Altman wrote that the autoswitcher broke and was out of commission for a chunk of the day, and the result was GPT-5 seemed way dumber. Users who had asked ordinary questions had been served the fast model for everything, including questions the fast model could not handle. The same day, legacy models had been removed from the picker for non-Pro users without warning, so there was no way to compare or fall back. Complaints about tone arrived alongside the complaints about capability. People described the new model as flat and lobotomised next to GPT-4o.

The response came fast. Altman said GPT-5 would seem smarter starting that day, committed to restoring GPT-4o as an option for Plus users, and by August 13 said the team was working on making the model feel warmer. A personality update rolled out on August 15. The weekly limit for GPT-5 Thinking for paying users settled at 3,000 messages, a number that only makes sense if users are choosing the reasoning model directly rather than leaving the choice to a router.

Routing as an answer to the compute question

Strip away the launch and the design is sound. The cost of a reasoning model is dominated by the thinking tokens, and most chat traffic does not need them. A classifier that sends five percent of queries to the expensive path and 95 percent to the cheap one is worth more, in serving cost, than most architecture improvements. It also changes the product question from which model should we pick to how much is this question worth, which is the question a good assistant should be asking anyway.

The catch is that the router is a model making a prediction about difficulty from the surface of a message, and difficulty is not on the surface. A one-line question about a drug interaction and a one-line question about the weather look alike to a classifier trained on conversation type. Explicit intent, which the system card lists as a routing signal, is the escape hatch. If you say think hard about this, the router listens. That means the user is still turning the knob. They are just doing it in prose instead of a menu.

What the rollback says about trust

We read the reaction as being about legibility rather than capability. When a user picks a model, a bad answer is the model's fault and the user knows which model to blame. When a router picks, a bad answer could be the fast model, the reasoning model, or the router, and the user cannot tell which. On launch day the failure was the router, and it looked to everyone like the whole system had regressed. A system that budgets its own effort has to make the budget visible, or every cheap answer becomes evidence that the product got worse.

The GPT-4o restoration is the sharper signal. People asked for the old model back rather than for a better router. Part of that was tone, and OpenAI treated it as tone. Part of it was that GPT-4o had a fixed, known behaviour, and a fixed behaviour is easier to build habits around than a variable one, even when the variable one is better on average. The 3,000 message limit on Thinking is OpenAI conceding the point. Users wanted a manual override, so the product now has one, with a quota attached.

What we would measure

The system card does not publish the router's error rates, and we do not expect it to. What we would want is a study on the API, where the reasoning level is explicit, that takes a sample of real chat prompts, runs each at minimal and at high, and records the fraction where the answers differ materially. That fraction is the router's job. If it is small, the router barely matters and the launch backlash was about the picker and the personality. If it is large, then routing accuracy is the capability and it should be evaluated and reported like any other.

Our guess is that the routers will get good enough that the toggle survives as a comfort rather than a necessity, the way a manual transmission survives. But that guess rests on the router becoming legible, and this week showed what happens when a system spends its own effort without telling anyone how much it spent.

Sources

  1. OpenAI: GPT-5 system card
  2. Wikipedia: GPT-5
  3. Simon Willison: GPT-5, key characteristics, pricing and model card