What shipped

OpenAI released two open weight models on August 5 under the Apache 2.0 licence, with the weights, inference code, tool environments and tokenizer all included. The model card gives the real sizes. gpt-oss-120b has 36 layers, 116.8 billion total parameters and 5.1 billion active per token. gpt-oss-20b has 24 layers, 20.9 billion total and 3.6 billion active. Both are mixture of experts transformers descended from the GPT-2 and GPT-3 architectures, with a residual width of 2880, RMSNorm and pre-LN placement.

The MoE blocks are where the parameters live. The 120b model has 128 experts per block and the 20b has 32, and both route each token to the top 4 experts, weighting the outputs by a softmax over only the selected experts. The card's parameter table makes the imbalance explicit. In the 120b model, 114.71 billion parameters are in the MLP blocks, 0.96 billion in attention and 1.16 billion in embeddings. The experts are more than 90 percent of the model.

The quantization

That imbalance is the whole story of the memory footprint. The model card says the MoE weights were post-trained with quantization to the MXFP4 format, at 4.25 bits per parameter, and that because those weights are over 90 percent of the parameter count, quantizing them lets the larger model fit on a single 80GB GPU and the smaller one run on systems with as little as 16GB of memory. The checkpoints come out at 60.8GiB and 12.8GiB. Willison ran the 20b model on a Mac and reported it using around 12 to 13GB of RAM.

The 4.25 figure is what MXFP4 costs. The format stores each weight as a 4 bit floating point value, and the extra quarter of a bit per weight is the amortised cost of the scale factors the format shares across groups of weights. The attention weights, embeddings and everything outside the experts are not quantized this way, and they are small enough that it does not matter. Four bits on 90 percent of the model and higher precision on the rest gives you almost all the compression at almost none of the risk.

Why 4 bit expert weights cost so little

The intuition for why this works is about which weights each token actually touches. In a dense model every parameter contributes to every token, so precision loss anywhere accumulates everywhere. In a top 4 of 128 routing scheme, a given token passes through a small fraction of the expert weights and the rest are silent. Quantization error in an expert only shows up on the tokens routed to it, and the attention and embedding weights that every token passes through are left at higher precision.

The other reason is in the phrase post-trained with quantization. The card does not describe the models as quantized after training with a calibration set, which is what most open 4 bit checkpoints are. It describes quantization as part of post-training, which means the model was trained through the low precision representation of its expert weights and had the chance to adapt to it. The MXFP4 weights are the weights the card evaluates. The card reports 2.1 million H100 hours for training the 120b model, with the 20b needing almost ten times fewer, and the quantization is inside that budget rather than a step somebody ran afterwards.

The rest of the design

A few architecture choices deserve a mention for anyone reading the card. Attention alternates between banded windows of 128 tokens and fully dense layers, following GPT-3. Each layer has 64 query heads of dimension 64 with grouped query attention over 8 key value heads. Rotary embeddings with YaRN extend the dense layers to 131,072 tokens of context. Each attention head has a learned bias in the softmax denominator, which the card compares to attention sinks and which lets a head attend to nothing at all.

The models come with three reasoning effort levels, low, medium and high, and the card reports benchmark results across all three. Willison timed the thinking on one prompt at 0.07 seconds at low, 4.44 seconds at medium and over five minutes at high. On GPQA Diamond the 120b model scored 80.1 percent against 81.4 for o4-mini, and the 20b scored 71.5. The knowledge cutoff is June 2024. Chat is handled through a format OpenAI calls harmony, which the models were trained on and which any serving stack has to reproduce exactly.

What changes if this becomes the norm

The important precedent is a frontier lab treating the quantized checkpoint as the release artefact rather than as a community afterthought. For the last two years the pattern has been that a lab ships bf16 weights and the ecosystem produces GPTQ, AWQ and GGUF variants of varying quality, each with its own accuracy loss that nobody measures carefully. If the lab does the quantization inside post-training, the accuracy loss is accounted for in the model's own evals and there is one canonical set of numbers.

What we want to see is a controlled comparison that the card does not give. Take the same recipe, run post-training with and without the MXFP4 constraint on the experts, and report the gap. The card's argument that the experts tolerate 4 bits is convincing on the reasoning and the benchmarks look fine, but the number we actually care about is how much a quantization aware post-training recovers relative to quantizing the same model afterwards. With Apache 2.0 weights and a published architecture, that is an experiment an academic group with a few hundred GPU hours could run on the 20b model, and it would settle whether this approach should be the default for every MoE release.

Sources

  1. OpenAI: gpt-oss-120b and gpt-oss-20b Model Card (arXiv PDF)
  2. Simon Willison: OpenAI gpt-oss