The problem with open VLMs

Most open vision-language models of the last eighteen months share a secret. Their training data was generated by a closed model. The standard recipe takes an image, sends it to GPT-4V with a prompt asking for a detailed caption, and trains on the output. The open model is then a distillation of the closed one, and as the Molmo paper puts it, the scientific community is missing foundational knowledge about how to build performant VLMs from scratch. You cannot learn how to make a capability if you copied it.

Matt Deitke, Christopher Clark, Sangho Lee, and about fifty co-authors at the Allen Institute for AI posted Molmo and PixMo on September 25. The models are open. The bigger claim is that the data is open too, and was collected without any external VLM in the loop.

Captions by voice

The core dataset is PixMo-Cap, 712k distinct images with 1.3M transcripts and captions. The collection method is the part we found clever. Annotators were asked to describe an image out loud for 60 to 90 seconds, and the audio was transcribed. Typing a detailed caption is slow and people stop early. Talking for a minute produces an average of 196 words per description, far longer than existing caption datasets, with the kind of spatial and relational detail that people say naturally and rarely bother to type.

The audio also serves a second purpose. It is proof that a person produced the description. A text caption could have been pasted from a model. A minute of someone talking about an image is much harder to fake at scale, and the team keeps the recordings as evidence that the data is what they say it is. That is a provenance argument built into the collection protocol.

The other datasets follow the same principle. PixMo-AskModelAnything has 162k question-answer pairs over 73k images, built by having annotators ask questions and refine answers with a text-only language model that never sees the image. PixMo-Points has 2.3M question-point pairs over 223k images, where the answer is a location in the image, which lets the model point and count by pointing. There are smaller synthetic sets for documents, clocks, and counting, generated by code rather than by a VLM.

The results

The models come in four sizes: MolmoE-1B, Molmo-7B-O on the OLMo base, Molmo-7B-D on Qwen2, and Molmo-72B on Qwen2 72B, all with a CLIP vision encoder. On an eleven-benchmark average the 72B model scores 81.2, the 7B-D scores 77.3, and the 1B mixture-of-experts model scores 68.6. The paper reports that the 72B outperforms Gemini 1.5 Pro, Gemini 1.5 Flash, and Claude 3.5 Sonnet on that average.

The human evaluation is the number we trust more, because benchmark averages over eleven sets reward whatever the sets happen to measure. In pairwise human preference the 72B model reaches an Elo of 1077, second only to GPT-4o among the models compared. An open model with open data being the second most preferred VLM in a human study is the result the paper should be remembered for.

Why refusing to distill was two decisions

The licensing decision is the obvious one. Captions generated by a commercial model come with that model's terms of service attached, and those terms typically forbid using outputs to train competitors. A dataset built that way is open in the sense that you can download it and closed in the sense that a lawyer would hesitate to let you ship a product on it. Voice captions from paid annotators have no such shadow. AI2 can release them and mean it.

The scientific decision is less obvious and we think more important. If your training data is a closed model's output, the ceiling of your model is the closed model, and the interesting question of what data actually teaches a VLM to see is unanswerable, because the answer is always whatever GPT-4V put in the caption. PixMo lets you ask the question. The paper's own conclusion is that the quality of the newly collected data was the most critical factor in the result, which is a claim you can only make when you know where the data came from.

What we would want checked

The result rests on the claim that no external VLM touched the data, and the audio recordings are the evidence. We would like an outside group to sample the recordings against the captions and confirm the match, because the whole argument depends on it and the paper is the only source so far. We would also like to see the ablation the paper invites: train the same 7B architecture on an equal number of GPT-4V captions and on PixMo-Cap, and report which one wins the human evaluation.

If PixMo wins, the field has been distilling from closed models out of habit rather than necessity, and the cost of a minute of a person's speech per image is cheaper than it looked. If the distilled model wins, AI2 has still given everyone a clean baseline, and the size of the gap tells us exactly what the closed models know that a person describing a picture does not.

Sources

  1. Deitke et al., Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models (arXiv 2409.17146)