Qwen3 and the hybrid thinking switch
Alibaba's Qwen3 family puts a reasoning mode and a direct-answer mode inside one set of weights, with a budget the caller controls, across eight models from 0.6B to a 235B mixture of experts. An explainer on how a single model is trained to think on demand, and what the design trades away.
What shipped on April 28
The Qwen team released eight models on Monday, all under Apache 2.0. Two are mixtures of experts: Qwen3-235B-A22B, with 235 billion total parameters and 22 billion active per token, and Qwen3-30B-A3B, with 30 billion total and 3 billion active. Six are dense: 32B, 14B, 8B, 4B, 1.7B and 0.6B. The dense models support context up to 128K tokens, with the two smallest at 32K. The flagship MoE was not yet available for public download at launch, according to TechCrunch, while the 32B dense model was.
The design decision that makes the release interesting is that every one of these models can operate in two modes. In thinking mode the model produces an extended chain of reasoning before its answer, in the style of QwQ or DeepSeek R1. In non-thinking mode it answers directly, like a conventional chat model. Until now those were two different products from most labs, with two different weights and two different prices. Qwen3 puts both in the same checkpoint and lets the caller choose per request.
How the switch works from the outside
There are two controls. The hard switch is an enable_thinking flag in the chat template, set true or false when the request is built. The soft switch is a pair of tags, /think and /no_think, which a user can append to any message in a multi-turn conversation to change the mode for that turn. The blog post gives a worked example of a conversation that flips between the two across turns, with the model following the most recent instruction.
On top of the binary switch there is a thinking budget, a cap on how many tokens the model may spend reasoning before it must answer. The Qwen post reports what it calls scalable and smooth improvements in performance as the budget increases, with a figure showing benchmark scores rising with allocated reasoning tokens. That is the property that matters for deployment. A caller with a latency constraint can set a budget rather than choosing between a fast wrong answer and a slow right one, and the trade-off is continuous rather than a cliff.
How the switch is trained in
The post describes a four-stage post-training pipeline, and the order is the whole story. Stage one is a long chain-of-thought cold start, fine-tuning on reasoning traces across maths, code, logic and STEM so the model learns the format. Stage two is reinforcement learning on reasoning tasks with rule-based rewards, scaling up the compute spent on exploration. At the end of stage two you have a reasoning model and nothing else, roughly where QwQ-32B was.
Stage three is the part they call thinking mode fusion. The reasoning model from stage two is fine-tuned on a mixture of long chain-of-thought data and ordinary instruction-tuning data, with the mode signalled in the prompt. The model learns that some requests get a reasoning block and some do not, and which is which. Stage four is general RL across more than twenty task types covering instruction following, format adherence and agent use, intended to fix behaviours that the earlier stages left rough without undoing them.
The choice to build the reasoner first and blend the direct-answer behaviour in afterwards, rather than the reverse, is worth dwelling on. It suggests the team found it easier to teach a strong reasoner to sometimes skip reasoning than to teach a chat model to reason. The smaller models in the family were trained with strong-to-weak distillation from the larger ones rather than running the full pipeline, which is how a 0.6B model ends up with a thinking mode at all.
What the pretraining had to supply
None of the post-training would work on a weak base. Qwen3 was pretrained on approximately 36 trillion tokens covering 119 languages and dialects, close to double the 18 trillion used for Qwen2.5. The corpus includes text extracted from documents with Qwen2.5-VL and synthetic maths and code generated with the Qwen2.5 maths and coder models. Pretraining ran in three stages: over 30 trillion tokens at 4K context to build basic skills, a further 5 trillion weighted toward STEM, code and reasoning, and a final long-context stage extending to 32K.
The size-for-size claims are aggressive. The post says Qwen3-4B can rival Qwen2.5-72B-Instruct, and that the 30B-A3B MoE beats QwQ-32B while activating a tenth of the parameters. TechCrunch reports the flagship's claims of beating o3-mini on Codeforces, AIME and BFCL. We have not run any of these ourselves yet, and reported numbers from a model's own release post are the ones most in need of a third-party check, so we are treating them as claims until the usual leaderboards catch up.
What the design trades away
Two costs come with a single hybrid checkpoint. The first is that the non-thinking mode is a fine-tuned behaviour laid over a reasoning model, and stage three can only preserve so much. If the fusion step degrades the reasoning ceiling relative to a pure reasoner, the family pays that tax in every deployment. The post does not report a stage-two-only ablation against the final model, and that is the number we would most like to see.
The second is that the switch lives in the prompt. A /no_think tag in user text changes model behaviour, which means anyone who controls part of the context, including a retrieved document, can in principle flip the mode. For a chat assistant that is a nuisance. For an agent that reads untrusted input and is expected to reason carefully before acting, it is a control surface that needs to be locked to the hard flag and filtered from content.
What we would want to try first is the budget curve on a task with a known difficulty distribution. If performance really is smooth in the budget, then per-request budget selection becomes a routing problem, and a small classifier that predicts how many reasoning tokens a query needs could recover most of the benefit of a large reasoner at a fraction of the average cost. That experiment is now cheap to run, because the 4B and 8B models fit on one consumer GPU and the licence permits anything.
Sources
From the foundation