Grok-1 by torrent: 314 billion parameters, Apache 2.0, no paper
xAI put a 314 billion parameter base checkpoint on a magnet link under the most permissive licence going. What it left out shows the gap between open weights and open science.
What was released
On Sunday xAI published the weights of Grok-1, the base model behind its Grok assistant, under the Apache 2.0 licence. The announcement is short. The model has 314 billion parameters in a mixture-of-experts arrangement with 25 percent of the weights active per token. What is on offer is the raw base checkpoint from the pretraining phase, which the post says concluded in October 2023. It has not been fine-tuned for dialogue or anything else.
The GitHub README fills in the architecture. Eight experts with two used per token. 64 layers. 48 query heads and 8 key-value heads. Embedding size 6,144. A SentencePiece tokeniser with 131,072 tokens. Rotary position embeddings, activation sharding and 8-bit quantisation support, and a maximum context of 8,192 tokens. The training stack is described in one phrase, JAX and Rust. The checkpoint ships as int8 and the example code notes that a multi-GPU machine is required to run it.
Distribution is by magnet link and, since yesterday, through the Hugging Face hub. The choice of a torrent for the first release is unusual for a company and a reasonable engineering decision for a file of this size. It also set the tone. This is a drop rather than a publication.
What was not released
The list is longer than the release. There is no description of the training data beyond the phrase a large amount of text data. There is no data card, no filtering description, no deduplication note, no statement about what was excluded. There are no evaluation results, on any benchmark, in either the announcement or the README. There is no training recipe, no optimiser configuration, no schedule, no compute figure, no loss curve. There is no safety or alignment section because there is nothing to align. It is a base model.
There is also no technical report. For a model in this parameter class that is the notable absence. Every other large release of the past year has come with at least a page of numbers, and most have come with a document that lets a reader form a view on whether the model is worth the cost of running it. Grok-1 arrives with a licence and a shape.
Open weights versus open science
Apache 2.0 is about as open as a licence gets. You can run the model, fine-tune it, sell what you build on it, and never tell xAI. On the licensing axis this release is more open than Llama 2, which carries use restrictions, and more open than any lab's flagship. If open meant only the licence, this would be the most open large model in existence.
It is much harder to do science with. To reproduce a result you need the data or a description of it. To compare fairly you need evaluations run under stated conditions. To understand a behaviour you need to know what the model was trained to do. None of that is here. What you can do with Grok-1 is measure it yourself, which is a real contribution, and it is a contribution that puts all the cost of understanding the artefact on the people who receive it.
The comparison we keep coming back to is the small open models with full recipes. A model an order of magnitude smaller, released with its data, its code and its logs, tells the field more about how to train language models than this checkpoint does. Grok-1 tells us that a 314 billion parameter MoE with these dimensions can be trained. It does not tell us how, or how well.
Why release it this way
We can think of three honest reasons. The first is that the model is superseded. Pretraining ended in October and xAI has presumably moved on, so the marginal cost of releasing it is low and the marginal cost of documenting it is high. The second is that documentation invites scrutiny of the data, and no lab currently wants to describe its book corpus in writing. The third is that a torrent and a licence generate goodwill on their own, and a table of middling benchmark numbers would subtract from it.
None of those is a criticism of the people who did the release. Shipping a checkpoint this size cleanly, with working inference code and a permissive licence, is real work. It is a criticism of the word open when it is used to describe the result. Weights without provenance are an artefact. Science is the record of how the artefact was made.
What we would like someone to do with it
The useful project is the one xAI did not do. Run the standard suite on the base checkpoint under documented conditions and publish the table. Probe the tokeniser and the model's behaviour on known benchmark items for signs of contamination, which is the only handle we have on the data. Fine-tune a small instruction variant and report how far it lands from Grok-1 as deployed. That would turn the drop into a data point.
And for whoever writes the next release post of this kind, the ask is modest. One page. What the data was in broad strokes, what the model scores on five common benchmarks, and how much compute it took. That is the difference between handing the field a file and handing it a result.
Sources
From the foundation