The question the post answers

Low rank adaptation has been the default way to fine-tune a large model on a small budget since 2021, and the standing worry about it has always been the same. It trains a small number of parameters, so it must be leaving something on the table compared with updating every weight. The usual practice has been to use LoRA when you cannot afford full fine-tuning and to feel slightly guilty about it.

John Schulman's post at Thinking Machines, published on September 29, sets out to measure the gap directly and finds that under the right settings there is not one. The title is the claim. If you configure LoRA properly, you do not regret not having done full fine-tuning. The post is unusual in that it is mostly plots of learning curves with matched compute, and the conclusions are the kind you can check.

The three settings that matter

The first finding is about learning rate. The optimal learning rate for full fine-tuning is about ten times lower than the optimal rate for a high rank LoRA, and for short runs of around a hundred steps the ratio rises to roughly fifteen. Most of the historical underperformance of LoRA, on this reading, came from people transferring a full fine-tuning learning rate across and training the adapter too gently. The post also observes that the 1/r scaling factor built into LoRA makes the optimal rate roughly independent of rank, so a rate tuned at rank 8 transfers to rank 256.

The second is about where the adapter goes. Applying LoRA only to attention projections, which was the original paper's recipe and remains common, underperforms badly. In their comparison an attention-only adapter at rank 256, about a quarter of a billion parameters, lost to an MLP-only adapter at rank 128 with a similar parameter count. The parameter budget is not the issue. The MLP layers hold most of the capacity in a transformer, and an adapter that skips them is fighting with one hand. The recommendation is to put LoRA on every matrix.

The third is about rank, and it depends on what you are training. For supervised fine-tuning on large datasets, low ranks eventually fall behind full fine-tuning once the dataset contains more information than the adapter can absorb, which is what you would expect. For reinforcement learning the picture is different. LoRA at rank 1 matched full fine-tuning on their reasoning tasks. The post's explanation is an information argument. A policy gradient step carries about one bit per episode, so an RL run on mathematical reasoning needs only on the order of 320,000 bits of capacity, and a rank 1 adapter on a large model already has around three million parameters.

The caveats the post includes

Two costs are reported rather than hidden. LoRA is more sensitive to batch size. As the batch grows past some point, LoRA takes a larger penalty in loss than full fine-tuning does, and the effect is independent of rank, which the authors attribute to the product-of-matrices parametrisation rather than to capacity. If your setup relies on very large batches for throughput, the equivalence weakens.

The other is the compute accounting. A LoRA training pass costs slightly more than two thirds of the FLOPs of a full fine-tuning pass, because the backward pass through the frozen weights still has to happen. LoRA saves memory and optimizer state far more than it saves arithmetic. The advantage the post claims is that with matched sample efficiency you get that memory saving for free. Training does not become dramatically cheaper per step.

Tinker as the product form of the argument

On October 1 the same lab announced Tinker, a hosted fine-tuning API. The interface is two primitives, forward_backward and sample, and the user writes the training loop in Python while the service handles distributed execution on open weight models, including mixture of experts models like Qwen-235B-A22B. Switching model size is a one string change. The service is free to start with usage pricing to follow.

The connection to the post is stated in the announcement. Tinker uses LoRA so that many training runs can share the same base model on the same hardware, which is what makes a hosted service economical. That only works as a product if LoRA is not a compromise, which is exactly what the post two days earlier set out to establish. Read together, the research post is the justification and the API is the bet. Beta users named in the announcement include the Goedel theorem proving team at Princeton, a chemistry group at Stanford, Berkeley's SkyRL group doing asynchronous off-policy RL with tool use, and Redwood Research running RL on AI control tasks.

What this means for a team deciding whether to fine-tune

The question has changed shape. A year ago the choice was between prompting, LoRA as the cheap and slightly worse option, and full fine-tuning as the expensive and correct one. If the findings hold, the middle option is the correct one for almost every case that fits in the adapter's capacity, and for RL that is nearly every case. The cost of a serious fine-tuning experiment then falls to whatever a hosted service charges for forward and backward passes, plus the effort of writing the loop.

The recipe to copy is short. Adapter on every matrix. Learning rate about ten times what you would use for full fine-tuning. Rank chosen by how much information the dataset carries, which for RL means it barely matters. Watch the batch size. And measure against a full fine-tuning baseline at least once on your own task, because the post's curves are on their tasks and the equivalence is an empirical result with stated conditions rather than a theorem.

What we want someone to test is the rank 1 RL result on a harder domain than mathematical reasoning, where the reward is noisier and the episodes longer. If a one bit per episode budget still fits in a rank 1 adapter on, say, multi-step tool use with sparse success, then the capacity argument generalises and the decision about whether to fine-tune becomes purely a question of whether you have a reward signal. If it does not, we will learn where the information argument breaks.

Sources

  1. John Schulman and colleagues, LoRA Without Regret (Thinking Machines Lab, September 29, 2025)
  2. Thinking Machines Lab, Announcing Tinker (October 1, 2025)