Open weights were never the whole story

Most of what gets called an open model is a set of weights and a licence. The training code is sometimes published, the data mixture is described in a paper, and the hundreds of decisions in between, the runs that were abandoned, the learning rate that was changed at step 40,000, the data source that was dropped after it hurt an eval, live in a lab's private chat. You can use the model. You cannot see how it was made, and you cannot see what was tried and rejected.

Marin, announced on 19 May by David Hall and colleagues at Stanford's Center for Research on Foundation Models, with Percy Liang as senior advisor, is an attempt to publish that middle part. The project calls itself an open lab. The claim is that development is visible from the start rather than disclosed at the end, and that the artefact being released is the process as much as the model.

How an experiment moves through the repo

The workflow is borrowed from software engineering and from preregistration. An experiment begins as a GitHub issue that states the hypothesis and the intended measurement before anything runs. The experiment itself is written as Python code that specifies the data, the model configuration and the evaluation, and it arrives as a pull request. Other participants review it the way they would review code, with the announcement drawing an explicit comparison to open peer review on OpenReview.

When the run executes, its progress is visible live through Weights and Biases, and the results are posted back to the originating issue automatically. That closes the loop. The hypothesis, the code, the review comments, the training curves and the outcome sit in one thread with a permanent URL. A failed experiment leaves the same trail as a successful one, which is the part most labs never publish.

The models that came out of it

The first release is Marin 8B Base, a Llama-style dense transformer trained on 12.7 trillion tokens, in a run the team named deeper-starling. The announcement reports that it beats Llama 3.1 8B Base on 14 of 19 standard base model evaluations, and immediately qualifies the number with warnings about possible train-test overlap and sensitivity to prompting. We appreciate the caveat being in the same paragraph as the claim, since it usually lives in an appendix.

Marin 8B Instruct was produced by supervised fine-tuning on roughly 5 billion tokens. It does better than OLMo 2 on instruction following and worse than Llama 3.1 Tulu, and the team attributes the gap to not having done any RLHF or DPO stage yet. Together AI hosts the instruct model for people who want to try it without downloading weights. The compute for all of this came almost entirely from Google's TPU Research Cloud, with training built on JAX and Ray through the Levanter stack.

Two ways to contribute without a cluster

The announcement is honest that most people cannot run a 12.7 trillion token experiment, so it offers two smaller doors. The speedrun leaderboard, modelled on the nanoGPT speedrun, accepts optimisation submissions at several fixed compute budgets, so an algorithmic idea can be tested and ranked at a scale a university group can afford. The second is Datashop, which lets a domain expert describe the data they want in a prompt, has a classification pipeline pull matching material from a large corpus, and feeds the result into fine-tuning. The expert never touches the training infrastructure.

Both mechanisms are ways of decomposing a lab into contributions that fit in a pull request. That is the real design bet. Whether they attract enough contributors to matter is something the issue tracker will show in a year.

Does the workflow scale beyond a university

The obvious objection is that this works because nobody at Marin is trying to win. A commercial lab will not preregister its next data mixture in public, and it will not want competitors reading the review thread where its pretraining recipe was argued over. We think that objection is right and also beside the point. The value of Marin is that it produces a public record of what a careful 8B training run looks like in 2025, and that record is useful to everyone regardless of whether the big labs adopt the practice.

The less obvious objection is about review bandwidth. Software projects work with pull requests because a reviewer can read a diff in minutes. Reviewing a proposed training run means judging a data mixture, a hyperparameter choice and an evaluation plan, and the number of people qualified to do that is small. If the queue backs up, the preregistration discipline will erode and the issue tracker will become a log rather than a review.

What we would want to see next

The test of an open lab is what it does with a failure. We would like to see a negative result that changed the recipe, documented from hypothesis through to the decision, with the review comments intact. Marin's structure makes that possible in a way a paper never has. Whether the community writes it down when it happens will tell us whether open development is a new practice or a nicer way of publishing a model card.

The other thing we want is for someone to fork it. The code, the data recipes and the workflow are all public. A second group running the same recipe on different hardware, or running the same issue tracker for a different model family, would show whether the process transfers or whether it depends on this particular team.

Sources

  1. Marin, Announcing Marin: an open lab for building foundation models, 19 May 2025