Qwen 3.8 and the overthinking default
Qwen 3.8 27B ships with its reasoning effort set to the highest level, and on a simple drawing prompt it spent 22,000 thinking tokens and 21 minutes producing what it can produce in about two minutes with reasoning off. On why labs keep raising the default, and what it costs.
The release and the default
Alibaba's Qwen team released Qwen 3.8 27B on August 15 under Apache 2. It is a dense 27 billion parameter model with vision, a context window of 262,144 tokens, and a Q4_K_M quantisation that fits in 17 GB. Simon Willison ran it on a 128 GB M5 Max MacBook Pro and on an NVIDIA DGX Spark through LM Studio and llama-server, and his writeup is the most detailed hands-on account we have seen so far. Most of what follows draws on it.
The detail that made us want to write this is a configuration choice. The model exposes a reasoning effort setting with low, medium and xhigh levels, and it ships with xhigh as the default. Willison calls that a hilarious default and says plainly that it is not a good way to run the model, especially on consumer hardware. His recommendation is to ignore it and start at low or with reasoning off entirely.
What 22,000 thinking tokens bought
The test that exposes the cost is his standard one, asking for an SVG of a pelican riding a bicycle. With xhigh reasoning the model produced 22,276 reasoning tokens and 3,223 output tokens, and at the 15 to 30 tokens per second he was getting in LM Studio the whole thing took 21 minutes. With reasoning disabled the same prompt produced 3,715 output tokens in 137 seconds. The drawing was competent either way. Twenty one minutes versus a little over two, for roughly the same picture.
The circle test is the one we keep coming back to. He asked for an SVG of a circle. The reasoning trace begins by acknowledging that this is a simple request and then announces that the model wants to make it a carefully crafted piece anyway, and several minutes later it delivers an animated circle with gradients, rotating dashed rings and a pulsing glow. The output is lovely and it is the wrong answer. Someone who asks for a circle wants a circle.
This is what overthinking looks like from the outside. The model reaches the answer quickly and then refuses to accept that the answer is small, spending its budget elaborating a task that did not ask for elaboration. The bounding box test shows the same tendency in a different setting. Asked for JSON coordinates on a pelican photo, the model returned accurate boxes, and then its reasoning trace led it to build a full HTML visualisation tool with a demo scene of stylised pelicans that nobody had requested.
Why the default keeps going up
We can think of three reasons a lab would ship a model at maximum effort, and none of them is that users want it. The first is benchmarks. Reasoning effort is one of the strongest levers on evaluation scores, and a release is judged on its launch numbers. If the reported scores were obtained at xhigh, then shipping the default at low would mean that anyone who tries the model out of the box gets a worse model than the one in the table. Setting the default to match the benchmark condition makes the first impression consistent with the marketing.
The second is asymmetry in how failures are attributed. A model that thinks too long on an easy question annoys the user. A model that thinks too little on a hard question gets it wrong, and the wrong answer is screenshotted. From a lab's point of view the second failure is more expensive than the first, so the default drifts toward more thinking even when the median request does not need it.
The third is that the cost is externalised. On a hosted API the lab sees the extra tokens as revenue. On a local machine the user sees them as 21 minutes of fans. Willison's throughput of 15 to 30 tokens per second is a fraction of what hosted endpoints deliver, and he notes that OpenAI's models were running at 74 to 184 tokens per second in his comparisons.
What overthinking actually costs
The direct cost is time and tokens, and the pelican figures put it at roughly a tenfold increase in wall clock for no visible gain on that task. The indirect cost is harder to see and probably larger. A model that runs for 21 minutes does not get used for the small tasks that make up most of a working day, so the question of whether it is any good at those tasks never gets asked. Willison's own conclusion is that the model is a real general purpose assistant in 17 GB, which he calls a miracle, and that speed is the main thing keeping it from being a daily driver. The default makes that problem several times worse than the hardware requires.
There is also a quality cost when the elaboration replaces the answer. In an agent loop the same tendency turns into extra files, extra abstractions and features that were never in the ticket. Willison had the model drive a coding agent through the Pi framework, and it read a codebase and wrote and tested a working JSONL to Markdown converter without help, which is the good news. The bad news is that the same instinct that produces rotating dashed rings on a circle will produce a configuration system on a script.
Can a model budget its own thought
The setting exists because the model cannot yet decide for itself. A three-level knob is a confession that the question of how much to think has been handed back to the user, and the default is where the lab put it when it stopped deciding. What we want is a model that reads the request, estimates how much reasoning the request warrants, and spends that much. The circle trace shows the model already knows the request is simple. It says so in the first sentence. It then overrides its own judgement, which suggests the instinct to elaborate was trained in on top of a correct assessment.
That is a training target rather than a decoding trick. If the reward during reasoning training only credits correct answers, the model learns that more thinking is never punished, and the default drifts up in the weights the same way it drifts up in the configuration file. A reward that charges for tokens on tasks where the short answer was already right would push the other way. We would like to see a small open model trained both ways on the same data and then handed the circle prompt. Until then, the practical advice is Willison's. Turn the knob down first and only turn it up when the model gets something wrong.
Sources
From the foundation