Two releases in one week

On 5 June Google published quantisation-aware training checkpoints for the Gemma 4 family, which spans the E2B and E4B edge models, a 12B dense model and a 26B mixture-of-experts model. The checkpoints ship in two formats. One is the familiar Q4_0 layout that llama.cpp and its descendants have used for years. The other is a new mobile-specialised scheme, and it is the second one that moves the needle. In that format the E2B model fits in 1 GB of memory, and a text-only E2B without per-layer embeddings comes in under 1 GB.

The same week we got an independent project called TurboFieldfare running on the oldest laptop in our office. It is a Swift and Metal runtime, not built on any existing framework, that runs Gemma 4 26B-A4B, the instruction-tuned mixture-of-experts model with 26 billion total parameters and about 3.88 billion active per token. The author validated it on an 8 GB M2 MacBook Air with around 2 GB of weights plus KV cache resident. The full checkpoint on disk is 14.3 GB. Only a small slice of it needs to be in memory at any moment.

Two years ago running anything of this size locally on a base-model laptop meant swap, a fan at full speed and a token every few seconds. Both of these releases attack that from a different side, and together they mark the point where we stopped treating on-device inference as a demonstration.

What quantisation-aware training buys

Post-training quantisation takes a finished model and rounds its weights to fewer bits. It is cheap and it works reasonably well at four bits, but every rounding step is a small error the model never got to compensate for. Quantisation-aware training instead simulates the rounding during a final phase of training, so the weights settle into values that survive the rounding. Google reports that the QAT checkpoints reach higher overall quality than standard post-training quantisation baselines, though the blog post does not publish per-benchmark numbers, so we are reporting the claim rather than confirming it.

The mobile format is where the engineering is. It uses static activations, meaning the scaling factors are computed once ahead of time rather than per input, which suits mobile accelerators that dislike dynamic shapes. Weights are quantised channel-wise. Token-generation components, the decode path that runs once per output token, are targeted for 2-bit quantisation, which is where most of the memory saving comes from. Embeddings and the KV cache get their own treatment. The result is a model that runs across Hugging Face, llama.cpp, Ollama, LM Studio, vLLM, SGLang, MLX, Transformers.js, LiteRT-LM and Unsloth on the day of release.

How a 26B model fits in 2 GB

TurboFieldfare exploits the structure of the mixture-of-experts model. The 26B-A4B has a shared core that every token passes through and a set of routed experts of which only a few fire per token. The runtime keeps the shared core resident, about 1.35 GB in memory, and streams experts in from disk through a cache with configurable slots and least-frequently-used eviction. Quantisation is MLX affine 4-bit with group size 64 for the shared and routed experts, with an 8-bit router. The router is small and its precision matters, so it gets more bits.

The KV cache is stored in FP16 with a trick that follows the model architecture. Gemma 4 mixes sliding-window attention layers with full-attention layers. TurboFieldfare gives the 25 sliding-window layers a bounded circular buffer, since they never look further back than the window, and gives the 5 full-attention layers a linear buffer that grows with context. That keeps cache memory roughly proportional to the number of layers that actually need long context. The default cache size is 4K tokens.

The speeds are honest about the trade. On the 8 GB M2 Air the author measures 5.1 to 6.3 tokens per second. On an M5 Pro with 24 GB it is 31 to 35. The lower figure is usable for a chat and painful for anything that generates a long document. The runtime requires macOS 26, Metal 4 and Swift 6.2, so it is a current-generation tool, and the repository is Apache 2.0 with 38 commits at the time of writing.

Where on-device still loses

We tried the E2B QAT checkpoint and the 26B-A4B under TurboFieldfare on three internal tasks last week, and the picture is mixed in a way we want to describe precisely. For short structured extraction from documents that must not leave the machine, the E2B model in 1 GB is now good enough that we would not send the data anywhere. For summarising long transcripts, the 26B model at six tokens a second on an Air takes minutes per document, and a hosted model at a fraction of a cent per call finishes before the local one has loaded its experts.

The economics have a simple shape. On-device wins when the data is sensitive, the network is absent, the volume is high enough that per-call pricing adds up, or the latency of a first token matters more than the throughput of the rest. A cheap API call wins on long outputs, on anything that benefits from a model larger than what fits, and on the many cases where nobody has time to manage a runtime that needs a specific OS release. Neither of these has changed. What has changed is that the first list now includes real models rather than toy ones.

The next thing we want to see is a QAT checkpoint for the 26B model in the mobile format, since Google's announcement covers the family but the 1 GB headline is for E2B. If the 2-bit decode path holds quality at 26B, and if a runtime like TurboFieldfare can use it, the 2 GB figure becomes a 1 GB figure and the M2 Air numbers roughly double. That is the experiment we would run first if we had the checkpoint.

Sources

  1. Google: Quantization-aware training for Gemma 4
  2. TurboFieldfare on GitHub