Phi-3 and the model that runs locally on your phone
Microsoft's phi-3-mini has 3.8B parameters, quantises to about 1.8GB, runs at more than 12 tokens per second on an iPhone 14, and reports 69 percent on MMLU and 8.38 on MT-bench. A look at the curated-data recipe behind it and at the gap between those scores and what the paper itself admits the model cannot do.
The claim
The Phi-3 technical report describes a 3.8B parameter model, phi-3-mini, trained on 3.3T tokens, that scores 69 percent on MMLU and 8.38 on MT-bench. The authors say that puts it on par with Mixtral 8x7B and GPT-3.5, and they make the point with a photograph: the model quantised to 4 bits, occupying about 1.8GB, running natively and fully offline on an iPhone 14 with an A16 chip at more than 12 tokens per second. Mixtral, the comparison they lean on most, has 45B total parameters.
Two larger siblings are in the same report. Phi-3-small at 7B parameters trained on 4.8T tokens reaches 75 percent on MMLU and 8.7 on MT-bench. Phi-3-medium at 14B, same token count, reaches 78 percent and 8.9. Architecturally phi-3-mini is deliberately boring: a Llama-2 style decoder with the same tokeniser and a 32,064 vocabulary, 3,072 hidden dimension, 32 heads and 32 layers, trained in bfloat16. Every tool built for Llama 2 works on it unchanged, which is a good decision if your goal is deployment.
The recipe is the data
The paper is explicit that the model's size is not where the innovation is. The training data is heavily filtered public web data plus synthetic data generated by larger models, and pretraining runs in two sequential phases. Phase one is mostly web sources chosen to teach general knowledge and language. Phase two mixes an even more heavily filtered subset of that web data with synthetic data aimed at reasoning and niche skills.
The authors call this the data optimal regime, and they use the word optimal, as they say in a footnote, in an aspirational sense. The idea is that for a model this small you should calibrate the data to what the model has room to learn, rather than train on a fixed corpus and scale compute. Their Figure 3 plots the phi series against Llama-2 7B through 70B trained on the same data and shows the phi models above the curve. The comparison is fair only to the extent that the Llama-2 data is a reasonable baseline, and since the whole point is that the phi data is different, this is a plot that shows the intervention worked without isolating why.
What the paper says the model cannot do
The weakness section is short and we think honest. The model does not have the capacity to store much factual knowledge, and the authors point to low TriviaQA performance as the visible symptom. They suggest search augmentation as the remedy, which is a way of saying that the phone model is a reasoning and language engine that needs to be attached to something that knows things. The other admitted limitation is that the data is mostly English. The report also notes that factual inaccuracies and bias remain, as with most models, despite the responsible AI work.
Those two admissions explain the reaction we have seen from people who downloaded the weights this week. On the benchmark it looks like GPT-3.5. In conversation it feels like a much smaller model the moment you ask it something that depends on a fact it never had room to store. Both observations are correct, and the paper predicts the second one.
The benchmark-versus-vibes gap
There is a specific reason small models trained on curated data look better on tests than in use. MMLU is multiple choice over academic subjects and MT-bench is 80 two-turn questions graded by GPT-4. A data pipeline that filters the web for educational value and adds synthetic reasoning data is, by construction, a pipeline that produces text shaped like those evaluations. The authors report that all comparisons run through the same evaluation pipeline, so this is a distribution match rather than contamination in the narrow sense, a match between the training recipe and the tests, and it will make any benchmark that resembles the recipe an optimistic estimate of general use.
The way to read a paper like this is to trust the on-device numbers, the 1.8GB and the 12 tokens per second, because those are measured on hardware and hard to fake, and to treat the quality numbers as an upper bound that the TriviaQA line has already started to lower. What we would want next is a benchmark for small models that measures exactly what the recipe cannot teach: long-tail facts, non-English prompts, and questions phrased nothing like a textbook. If phi-3-mini holds up on that, the phone model story is real. If it does not, we have a very good demonstration of how much MMLU can be moved by choosing what to read.
Sources
From the foundation