What was in the box

On February 1 the Allen Institute for AI posted OLMo, a 7B model and a 1B model, and the paper's opening move is a list of what ships with them. The weights, under Apache 2.0. The full training data, Dolma, the same multi-source corpus the models were trained on. The training code, including the data ordering. Hundreds of intermediate checkpoints from the run. The evaluation code, Catwalk for downstream tasks and Paloma for perplexity across domains. The training logs. Inference code and an adaptation pipeline.

Each of those pieces had appeared before somewhere. The paper's own comparison makes the point. Pythia released checkpoints and data but sits well below current capability. Llama 2 released weights and a detailed report but no data. Mixtral released weights and little else. MPT and Falcon released weights and partial descriptions of data. What had not existed was a model close to Llama 2 in capability with every one of those artefacts open at once.

The models

OLMo-7B has 32 layers, a hidden size of 4,096 and 32 heads, and was trained on 2.46 trillion tokens from Dolma at a global batch of around 4 million tokens. OLMo-1B has 16 layers and was trained on 2 trillion tokens. The architecture follows the choices that Llama and PaLM made standard, with no bias terms, a non-parametric layer norm, SwiGLU activations and rotary embeddings, and the paper documents where it differs from each of the comparison models.

The 7B run happened on two clusters, LUMI with AMD MI250X GPUs and a MosaicML cluster with NVIDIA A100s, and the paper reports near-identical results from both, which is itself a useful data point for anyone worried about hardware lottery effects. Zero-shot on the eight core tasks, OLMo-7B averages 69.3 against 70.5 for Llama 2 7B, 70.3 for Falcon-7B, 69.8 for MPT-7B and 63.0 for Pythia 6.9B. It is behind the best of its comparison set by about a point and ahead of the only other 7B model with open checkpoints and data by six.

Why the data matters most

Of everything in the release, Dolma is the piece that changes what research is possible. With the data and the data order, you can ask which documents a checkpoint had seen when a given capability appeared. You can test memorisation against the actual training set rather than a guess at it. You can measure contamination of an evaluation directly. Every one of those questions is unanswerable for Llama 2, and the field has been answering them by proxy, running experiments on Pythia and hoping the results transfer to models ten times more capable.

The paper is honest about the trade-off this creates. On Paloma's per-domain perplexities OLMo-7B does best on data that resembles its own training mix and worse on domains, such as older web text, that Dolma deliberately excluded or decontaminated against. Falcon does best on RefinedWeb for the same reason. A fully open dataset makes that pattern visible. Closed datasets hide it without removing it.

Why the checkpoints matter almost as much

The intermediate checkpoints turn a model into a time series. The paper shows the eight core task accuracies rising over training, with a visible jump when the learning rate is decayed to zero over the last thousand steps, a trick the authors found by looking at exactly this curve. They compare the trajectory against Pythia-6.9B and RPJ-INCITE-7B, the other two models with public checkpoints, and note that the position of a checkpoint in the learning-rate schedule affects its evaluation score, which anyone comparing mid-run checkpoints across models should account for.

For interpretability work this is the difference between studying an organism and studying a fossil. Feature formation, circuit emergence, the point at which a model starts to do induction, all of these are questions about training dynamics, and they need checkpoints from a model whose data is also known. Pythia gave us that at small scale. OLMo gives it at a scale where the model is actually competitive.

What full openness costs and buys

The release has limits and the paper states them. The models are a point behind the strongest open weights at 7B. The data has been decontaminated against the evaluations it reports, which is correct practice but makes some comparisons with other models uneven. The paper also reports the carbon footprint of the run, which nearly no comparable release has done, and the two-cluster setup means the numbers are split.

What we take from the release is a standard. Before February it was possible to argue that a model at this capability could not be fully open, because nobody had done it. That argument is gone. A lab that releases weights without data and code is now making a choice, and it will have to explain the choice rather than claim it was unavoidable. What we would want next is for someone outside AI2 to reproduce a segment of the run from the released data and code, and report how close they get. Reproducibility is the claim the release makes, and it has not been tested yet.

Sources

  1. Groeneveld et al., OLMo: Accelerating the Science of Language Models (arXiv 2402.00838)