NeurIPS 2025 takeaways: 1000-layer RL, gated attention, and the hivemind
Notes from San Diego organised around the four best papers and three runners up. Depth scaling in self supervised RL, a sigmoid gate that removes attention sinks, a two timescale account of why diffusion models generalise, and a 26,000 query study of how alike language models have become.
Reading the awards as a signal
We came back from NeurIPS with a stack of notes and one overall impression. Three of the four best papers are about architecture and training dynamics in the plain sense, what happens when you change depth, change the attention nonlinearity, or look at the training curve of a generative model. After several years in which the biggest results came from scale and data, the awards committee chose papers that ask why a network does what it does. That reads to us as a field rediscovering architecture research, and we want to go through the papers in that light.
Depth as a capability axis in RL
The paper with the most surprising headline is from Kevin Wang, Ishaan Javali, Michal Bortkiewicz, Tomasz Trzcinski and Benjamin Eysenbach, titled 1000 Layer Networks for Self-Supervised RL. They train goal conditioned policies on locomotion and manipulation tasks with no rewards and no demonstrations, and scale the network to 1,024 layers. The reported gains over shallow baselines run from two times to fifty times depending on the task.
The claim that stuck with us is qualitative. At depth, the learned behaviours change character rather than just improving a score. Goal reaching that shallower networks never manage becomes reachable. Most RL scaling work has pushed on data, environments or model width. This paper says depth was an underused axis, and that the reason it was underused is that people assumed the instability at depth was fundamental rather than an engineering problem.
A sigmoid gate that removes attention sinks
Zihan Qiu, Zekun Wang, Bo Zheng and colleagues, in Gated Attention for Large Language Models, ran a systematic comparison of 30 attention variants. The winner is simple. After scaled dot product attention, apply a head specific sigmoid gate. That single change improved performance consistently, improved training stability, and improved scaling behaviour.
The explanation is that the gate adds nonlinearity and query dependent sparsity to a block that is otherwise linear in the value vectors. The consequence we found most useful is in the title. Models trained with the gate do not form attention sinks, the phenomenon where heads dump attention mass onto the first token because they have nowhere else to put it. A lot of long context engineering has worked around attention sinks. This paper suggests they are an artefact of a missing nonlinearity, and that you can train them away.
Why diffusion models do not memorise
Tony Bonnaire, Raphael Urfin, Giulio Biroli and Marc Mezard asked a question that has bothered us for a while. A diffusion model trained long enough on a finite dataset should eventually reproduce its training set, since that is the minimiser of the objective. In practice they generalise. The paper shows that training has two timescales. There is an early period during which the model learns the distribution and generalises, and a later period during which it begins to memorise. The gap between the two grows with dataset size.
That gap is the implicit regulariser. With enough data, the memorisation phase is so far out that ordinary training schedules never reach it, and you get generalisation for free without any explicit constraint. It is a clean dynamical story, and it also gives a concrete prediction, which is that on small datasets you should see memorisation appear at a predictable point in training. We would like to see the same analysis applied to language model pretraining, where the memorisation question is now a legal one as well as a scientific one.
The Artificial Hivemind
The fourth best paper is different in kind. Liwei Jiang, Yuanjun Chai, Margaret Li and coauthors built Infinity-Chat, a set of 26,000 real open ended user queries, sorted into six top level categories and 17 subcategories, with 31,250 human annotations and 25 independent annotators per example. Then they measured how alike model responses are.
They report two effects. Intra model repetition, where one model gives similar answers to the same open prompt across samples. And, more strongly, inter model homogeneity, where different models from different labs produce strikingly similar outputs to the same prompt. On open ended tasks where many good answers exist, the models converge on the same few. A second finding is that language models, reward models and LLM judges are all less well calibrated to human ratings on exactly the generations where human annotators disagree with each other, even when overall quality is comparable. The tools we use to score diversity are worst on the cases where diversity matters.
The authors frame this as a concern about the long term homogenisation of human thought. We would frame it more narrowly. If every lab trains on overlapping data, aligns with similar preference models, and evaluates with similar judges, convergence is the expected outcome, and this paper is the first large scale measurement of it.
The runners up and what to try
Three runners up are worth a line each. Yizhou Liu, Ziming Liu and Jeff Gore argue that representation superposition, encoding more features than dimensions, explains neural scaling laws, with loss scaling inversely with model dimension under strong superposition regardless of data distribution. Yang Yue, Zhiqi Chen, Rui Lu and colleagues show that reinforcement learning with verifiable rewards improves sampling efficiency but does not produce reasoning the base model could not already reach, since base models match or beat RLVR models at high sampling budgets. Zachary Chase, Steve Hanneke, Shay Moran and Jonathan Shafer settle a three decade old question with tight mistake bounds showing a quadratic advantage for transductive online learning.
The experiment we would most like to see is a combination of two of these. Take the gated attention change and measure whether it shifts the Hivemind homogeneity numbers at all. Our guess is no, because the convergence is coming from data and alignment rather than architecture. But it would be the first test of whether an architectural change moves an output diversity metric, and the awards this year suggest that is the kind of question the field is ready to ask again.
Sources
From the foundation