Apple Intelligence: a 3B model, many adapters, and a private cloud
Apple described a roughly 3 billion parameter on-device model, quantized to an average of 3.7 bits per weight, with rank 16 LoRA adapters swapped in per task and a server model behind Private Cloud Compute. Notes on what makes it the most concrete on-device architecture anyone has shipped.
What Apple actually said on Monday
On 10 June Apple announced Apple Intelligence and, on the same day, published a research post describing the models behind it. The consumer announcement was about Writing Tools, notification summaries, Genmoji and a Siri that can read the screen. The research post is the part we care about, because it is the first time a vendor has laid out, with numbers, how a language model is supposed to live on a phone.
The design has three pieces. A language model of roughly 3 billion parameters runs on the device. A larger model, size unstated, runs on Apple silicon servers under a system called Private Cloud Compute. And a set of small task-specific adapters is loaded on top of the on-device model depending on what the user is doing. The features ship in beta this autumn, in US English, on the iPhone 15 Pro and on iPads and Macs with an M1 or later.
The on-device model and its compression
The base model uses grouped-query attention and shares its input and output embedding tables, which saves parameters on a model where every megabyte counts. The on-device vocabulary is 49K tokens against 100K for the server model. Training ran on Apple's AXLearn framework, which sits on JAX and XLA, on a mix of licensed data and pages crawled by Applebot, with a filter pass to strip personal identifiers such as social security and credit card numbers, plus profanity and low-quality text. Apple says no private user data went in.
The compression is the interesting engineering. Weights are stored with what Apple calls low-bit palletization, a mixed 2-bit and 4-bit scheme that averages 3.7 bits per weight. They report being able to push to 3.5 bits without a significant quality loss. Bit rates per operation were chosen with an internal latency and power tool called Talaria. The quoted result on an iPhone 15 Pro is about 0.6 milliseconds of time-to-first-token per prompt token and a generation rate of 30 tokens per second, before any speculative decoding.
Those two latency figures translate into a product constraint. A 1,000 token email thread costs roughly 0.6 seconds before the first output token. A 100 word summary at 30 tokens per second takes a few seconds more. Both are within what a user tolerates for a button labelled Summarize, and neither would be if the model were twice the size at the same bit rate.
Adapters as the unit of capability
Rather than one model that does everything, Apple trains an adapter for each feature. The adapters are LoRA modules at rank 16, applied to the attention matrices, the attention projection and the fully connected layers. Adapter parameters are kept in 16-bit precision while the base stays quantized. For the 3B model each adapter is in the tens of megabytes, and Apple says they can be dynamically loaded, temporarily cached in memory and swapped as the user moves between tasks.
This matches what several of us have argued a device model should look like. One quantized trunk, which is the expensive artifact, and a library of cheap heads that can be retrained and shipped independently when a feature underperforms. A proofreading adapter and a notification summary adapter do not need to share a fine-tuning run, and a bug in one does not require re-releasing the base.
The evaluation story is thinner than the architecture story. Apple reports that the on-device model beats Phi-3-mini, Mistral-7B, Gemma-7B and Llama-3-8B on its own human evaluations, and that the server model compares favourably with DBRX-Instruct, Mixtral-8x22B, GPT-3.5 and Llama-3-70B. For summarization they sampled 750 responses per use case across email, messages and notifications and graded a response as good only if every dimension was good. On adversarial prompts they report the summarization adapter did not amplify sensitive content in over 99 percent of targeted examples. Those are Apple's graders on Apple's prompts, so we read them as evidence of internal diligence rather than as a leaderboard result.
Private Cloud Compute
When a request is too big for the device, it goes to Private Cloud Compute. Apple's security post sets out five requirements. Computation is stateless, so personal data sent to the server is not accessible to anyone other than the user and is not retained after the request. The guarantees are meant to be enforceable technically rather than by policy. There is no privileged runtime access, meaning no remote shell or interactive debugging on the nodes. Requests are non-targetable, so compromising one node should not let an attacker pick out a particular user. And the whole thing is meant to be verifiable by outside researchers.
The verification claim is the one to watch. Apple says production software images will be published within 90 days of deployment to an append-only transparency log, and that a device will only talk to a server whose image appears in that log. The nodes are custom Apple silicon servers with a Secure Enclave and Secure Boot, running a hardened operating system derived from iOS and macOS, with data volume encryption keys regenerated at every reboot. Researcher access and a virtual research environment are promised for later.
There is also a third tier. Some requests can be handed to ChatGPT, running GPT-4o, with the user asked for permission each time, IP addresses obscured, and OpenAI not storing the requests. That path is optional and separate from the Apple models, and Apple was careful to draw the line between it and Private Cloud Compute.
What we want to see next
The design is the most complete on-device recipe any vendor has published, and we expect other phone makers to copy its shape. What is missing is anything a third party can rerun. There is no model card with public benchmark scores, no released weights, and no adapter training code. The claim that a 3B model at 3.7 bits beats Llama-3-8B on useful tasks is plausible for narrow adapters and we would like to see it checked on a public instruction-following set.
The other open question is how many adapters a phone can hold before the swapping cost shows up, and whether users will feel the seam when a request is routed off the device. Apple has said nothing about the routing rule. That rule, more than the model, will decide how private the system feels in practice.
Sources
From the foundation