Is your provider serving the model you think? Kimi's vendor verifier
Moonshot published a test suite to check whether third-party hosts of its open weights behave like the reference deployment, after community benchmark reports kept coming back wrong. What the six checks catch, why open weights do not mean identical models, and what we would add.
The complaint that started it
Moonshot's account of why the verifier exists is short and a little embarrassing for everyone involved. After Kimi K2 Thinking shipped, the team kept getting community reports of benchmark scores that did not match what Moonshot measured. The first round of investigation found that most of the discrepancy was people running the model with the wrong decoding parameters, which is why the official API now enforces temperature 1.0 and top-p 0.95 in thinking mode rather than trusting the caller.
The second round was worse. When the team looked at third-party API deployments of the same weights, they found what they call stark performance gaps against the official endpoint. The line from the post that we keep coming back to is that the more open the weights are, and the more diverse the deployment channels become, the less controllable the quality becomes. That is a lab admitting that releasing weights means releasing control of what users experience under the model's name.
Why the same weights give different models
The post lists four failure sources, and each one is a place where a deployment can drift from the reference without anyone lying. Misused decoding parameters are the obvious one. KV cache bugs are subtler, since they only show up when a generation runs long enough for cached attention state to matter, and a short smoke test never triggers them. Quantisation degradation is the one people argue about most, because a host that serves an 8-bit or 4-bit version of a model to fit more replicas on a GPU will pass casual inspection and fail on the tails. And an incorrect chat template, meaning the wrapper that turns a message list into the token sequence the model was trained on, can quietly break tool calling while leaving plain chat looking fine.
None of these are exotic. They are the ordinary engineering differences between any two inference stacks. What the verifier does is treat them as testable claims rather than implementation details, and that reframing is the actual contribution.
What the six checks are for
The suite is built as a ladder, cheapest first. A pre-verification step confirms that parameter constraints are enforced before any benchmark is run, which catches the temperature problem at the door. OCRBench is a five-minute multimodal smoke test that fails fast if the vision pipeline is broken. MMMU Pro Vision stresses image preprocessing across more varied inputs. AIME 2025 is there as a long-output stress test, and the post says explicitly that its job is to expose KV cache bugs and quantisation loss, since a model that must produce tens of thousands of reasoning tokens will show both. The K2VV ToolCall benchmark measures whether function calls are triggered consistently and whether the arguments conform to the JSON schema. SWE-Bench sits at the top as a full agentic coding run, and that component is not open-sourced.
The reference numbers Moonshot publishes for its own endpoint give a sense of the settings: temperature 1.0, top-p 0.95, and max tokens of 16,384 for OCRBench, 65,536 for MMMU Pro Vision and 98,304 for AIME 2025, on which the reference scores 98.4 average over 32 samples. The tool-call test in the repository runs each case twice, once streaming and once not, reassembles streamed chunks before validation, and checks the returned arguments against the schema with a local JSON schema validator. Streaming is where template and chunking bugs hide, so running both modes is the right call.
The whole thing runs sequentially on two eight-GPU H20 servers in about 15 hours, with streaming inference, automatic retry and checkpoint resumption. The repository has since grown to cover later models and now publishes a table of providers, with Fireworks, Baseten, Together and a vLLM reference deployment scored alongside Moonshot's own endpoint on OCRBench, MMMU Pro Vision and two agentic benchmarks. The published numbers there sit within a few points of each other, which is what you would hope for after vendors have had the test to run against.
What the suite does not tell you
A verifier of this shape measures aggregate scores, and aggregate scores are exactly where quantisation hides. A 4-bit deployment can match the reference on AIME average while failing more often on the longest, hardest prompts, and a per-task average will not distinguish a host that is right 98 percent of the time for the right reasons from one that got lucky on the sample. The suite also tests against Moonshot's own endpoint as the reference, which means a bug in the reference implementation becomes the standard everyone else is graded against.
There is also an incentive question. Moonshot says it now works upstream with the vLLM and SGLang communities and validates vendors before release, which is the right place to fix things. But a vendor who knows the test can tune to it, and a suite whose hardest component is closed cannot be fully reproduced by the vendor being graded. We do not think that is bad faith. It is the same tension every benchmark has, moved from models to inference stacks.
What we would want next
The obvious extension is for the community to run this against every host that serves the weights, not only the ones that cooperated, and to publish the failures rather than the passes. The published table is a list of vendors who look fine. The interesting document is the one that says which hosts were serving a quantised model under the full model's name in the first weeks after release, and by how much tool-call accuracy dropped.
The larger point stands regardless of this one model. Open weights are a file. A model, as users experience it, is that file plus a template, a sampler, a cache implementation and a precision choice, and each of those is a decision someone else made on your behalf. Moonshot's closing line is that weights are open and the knowledge to run them correctly must be too. We would go further and say the tests should ship with the weights, so that anyone who downloads the file can also download the thing that tells them whether they are running it right.
Sources
From the foundation