The claim in three parts

On June 2 Peter Belcak, Greg Heinrich, Shizhe Diao, Yonggan Fu, Xin Dong, Saurav Muralidharan, Yingyan Celine Lin and Pavlo Molchanov posted a position paper from NVIDIA with a title that does its own arguing. Their definition of a small language model is functional. It fits on a common consumer device and runs inference at a latency that is practical for one user, which as of 2025 covers most models under 10 billion parameters.

The argument has three legs, labelled V1 to V3. Small models are already powerful enough for the language modelling errands inside agents. They are operationally a better fit, because agents call models through narrow interfaces for repetitive subtasks. And they are necessarily more economical, because the paper puts serving a 7 billion parameter model at 10 to 30 times cheaper than a 70 to 175 billion parameter one in latency, energy and FLOPs, with fine-tuning taking hours of GPU time rather than weeks.

The evidence they marshal

V1 is defended with a list. Phi-2 at 2.7 billion parameters runs about 15 times faster than 30 billion parameter models it is compared against. Phi-3 small at 7 billion matches 70 billion parameter models on code generation. SmolLM2, from 125 million to 1.7 billion parameters, is put level with 14 billion parameter contemporaries. Hymba-1.5B gets 3.5 times the token throughput of comparable transformers. DeepSeek-R1-Distill-Qwen-7B and Salesforce's xLAM-2-8B are cited as beating Claude 3.5 Sonnet and GPT-4o on their respective tasks.

Every one of those comparisons is on a benchmark chosen by the small model's authors, and the paper does not pretend otherwise. The more persuasive evidence is in the case studies. The authors went through three open source agent frameworks and estimated what fraction of their model calls could be handled by a specialised small model: around 60 percent for MetaGPT, 40 percent for Open Operator, and 70 percent for Cradle. Those numbers are judgment calls, and they are the kind of judgment call anyone running an agent in production could check against their own logs in an afternoon.

The conversion recipe

The practical heart of the paper is a six step algorithm for moving an agent from a frontier API to a set of small models. Log every call with encryption and access controls. Curate the logs and strip anything personal. Cluster the calls into task types without supervision. Pick a small model for each cluster based on capability and cost. Fine-tune it with a parameter efficient method such as LoRA. Then keep iterating as the distribution of calls drifts.

What we like about this recipe is that step one is free and step three is diagnostic. If you cluster your agent's calls and find that most of them are one of five things, the argument has already been won for your system. If you find a long tail of unique requests, it has been lost, and no amount of fine-tuning will save it. The paper's estimates suggest the first case is common, and we suspect they are right for agents that do one job.

Why it has not happened yet

The authors list three barriers. Large sunk investment in centralised inference infrastructure makes the marginal cost of one more frontier call look small. Small models are developed and evaluated against benchmarks designed for large general models rather than against agentic subtasks, so their fitness for the job is under measured. And small models get far less marketing attention, so the people choosing a model default to the name they know.

We would add a fourth that the paper touches on only through its heterogeneous systems argument. An agent that routes between a frontier model for conversation and small models for subtasks has two failure surfaces instead of one. Every small model is a place where the task distribution can drift out from under a fine-tune, and someone has to watch each of them. The 10 to 30 times saving is real per call, and the monitoring cost is real per model.

What we would want to see checked

The paper closes by inviting critique and promising to publish correspondence at a research.nvidia.com page. We would take them up on it with one request. Publish the call logs, or a synthetic proxy, for the three case studies, so that the 60, 40 and 70 percent figures can be reproduced by clustering rather than by reading. A position paper stands or falls on whether its central estimate survives contact with someone else's method.

The other test is time. If the argument is right, agent stacks built over the next year should show routing layers, task specific fine-tunes and a shrinking share of calls to the largest model. If they instead consolidate onto one frontier model with a bigger context window, the operational suitability claim was wrong about what builders value, even if the cost claim was right about what they pay.

Sources

  1. Belcak et al., Small Language Models are the Future of Agentic AI (arXiv 2506.02153)