The experiment

Jonas Geiping and Tom Goldstein posted a paper on December 28 that asks the question in reverse. Rather than how far scaling can go, how far can you get with one GPU in one day. The rules are strict. Any transformer trained from scratch with masked language modelling, on a single consumer card, in 24 hours of wall-clock time including data loading, followed by a short fixed fine-tuning protocol on GLUE that is excluded from the budget. They test a 2018-era RTX 2080 Ti and the more recent A4000 and A6000.

The result is that a crammed BERT reaches a GLUE score of 78.3 on the 2080 Ti and 78.6 on the A6000, against 80.9 for the original fully trained BERT-base checkpoint. Running BERT's normal training protocol for one day on the same hardware gets 52. The previous best attempt at fast BERT training, from Izsak and colleagues in 2021, gets 69.7 on the 2080 Ti. Two years ago that gap would have looked closed only by scaling up. The paper closes most of it by throughput.

Why architecture barely matters

The first section of results is the one we find most instructive, because it is mostly negative. They plot masked language modelling loss against tokens ingested for a range of architectural variants, and the curves lie almost on top of one another. Larger models learn more per token, but under a fixed time budget they process fewer tokens, and the two effects cancel almost exactly. Every reasonable architecture ends the day at an MLM loss around 1.9. The authors read this as scaling laws holding in the low-compute regime, and it means you cannot win by picking a cleverer shape.

What you can do is make each token cheaper without changing what the model learns. That is the lens for every architectural change they keep. Removing all the QKV biases and all the feed-forward biases speeds up gradient computation with no measurable effect on quality. Pre-normalisation helps, but mostly by stabilising training enough to allow a larger learning rate with less warmup. Swapping layer norm for RMS norm gives nothing. Adding a final layer norm and dropping the decoder bias help slightly. The tokeniser vocabulary matters up to a point and then plateaus. Sequence length is capped at 128, which is why attention cost is never the bottleneck here.

The training recipe

The learning rate schedule is tied to the budget rather than to a step count, so it decays as the day runs out. A one-cycle schedule with a peak of 0.001 gave the lowest loss, and the authors note that triangular schedules with fast annealing have better end-of-budget behaviour. The micro-batch that fits on a 2080 Ti is about 96 sequences, which is several times below the optimal batch, so they accumulate gradients and ramp the effective batch size aggressively over the run. Their optimum was around 1,536 for lowest pretraining loss but 4,032 for best downstream performance.

Dropout goes. In a single-epoch regime there is nothing to regularise against, and dropout reduces the effective number of gradient updates per second while barely reducing the runtime of each. They re-enable it at 0.1 only for fine-tuning. Sparse token prediction, computing the output layer only on masked positions, saves memory. All of these are throughput decisions. None of them change what the model is capable of learning, only how many tokens it sees before the clock stops.

Data was the surprising lever

The data section is where the numbers move most. Of four sources tested, a subset of the Pile was best out of the box, with MNLI accuracy of 80.5 against 79.8 for the classic BookCorpus and Wikipedia mix and 79.1 for a C4 subset. Exact substring deduplication, which helps at scale, did nothing here. What did help was filtering out sequences that the tokeniser found hard to compress, on the reasoning that such text is noise, and then sorting the training sequences by that same metric so the least likely ones are reached last or never.

On C4 the combination of filtering, sorting and a larger final batch took MNLI from 79.1 to 82.5, which is more than any architectural change in the paper delivered. Most of the gain came from C4 rather than BookCorpus and Wikipedia, which the authors attribute to C4 containing more of the noisy sequences the filter removes. The lesson we take is that at small scale, where every token has to count, the quality and ordering of the data is the one place you can still buy loss with engineering instead of compute.

What the paper is good for

Most pretraining papers change several things at once at a scale nobody can replicate. This one changes one thing at a time under a budget anyone with a workstation can match, and reports the ones that did nothing alongside the ones that worked. That makes it one of the few controlled ablation studies of the pipeline that exists, and we would trust its negative results more than most positive results from larger runs. The caveat is that the setting is narrow. Sequence length 128, encoder-only, masked language modelling, GLUE as the target. Whether the ranking of tricks survives a switch to decoder-only models at longer context is an open question.

What we would want someone to try is the same protocol with a causal language model and a 1,024-token context, on the same three GPUs. If bias removal, aggressive batch ramping and compressibility filtering still come out on top, we would have a small-scale recipe that transfers. If attention cost changes the answer, that would be worth knowing too.

Sources

  1. Geiping and Goldstein, Cramming: Training a Language Model on a Single GPU in One Day (arXiv 2212.14034)