The claim in one sentence

A trained language model cannot learn anything after training ends, except within the context window it is given, and Ali Behrouz and colleagues at Google Research think the reason is that we have been stacking the wrong thing. Their NeurIPS paper, titled Nested Learning: The Illusion of Deep Learning Architectures, argues that any model should be read as a set of nested, multi-level optimization problems, each with its own context flow, and that what we call depth is one particular and rather flat instance of that structure.

The provocation in the title is that architecture and optimizer are the same kind of object. The abstract states that known gradient-based optimizers, Adam and SGD with momentum among them, are associative memory modules that compress the gradients' information. We read the paper twice before that stopped sounding like a metaphor and started sounding like a definition.

Momentum as a memory

Here is the reading. An associative memory maps keys to values and is trained to reproduce that mapping. Momentum keeps a running state that is updated from each new gradient and read out to produce the parameter step. Nothing stops you from describing that state as a small memory whose keys are recent inputs and whose values are the local error signals, and whose update rule is itself a gradient step on a compression objective. Adam adds a second such memory for the squared gradient. The Google Research post says the same thing about backpropagation as a whole, that the training process can be modelled as an associative memory that learns to map data points to their local error values.

Once you accept that, the levels line up. The weights are a slow memory updated by the optimizer. The optimizer state is a faster memory updated by gradients. In a transformer the attention over the context is a memory that updates every token and forgets at the end of the sequence. Each level runs its own optimization over its own flow of context, at its own frequency. The authors call the abstraction that unifies these a continuum memory system, in which memory is a spectrum of modules each updating at a specific rate rather than a binary of short-term context and long-term weights.

Why depth was the wrong axis

The paper's argument about depth follows from that. Adding layers adds capacity at a single timescale, the one set by the outer optimizer. It does not add levels of learning. In-context learning, the paper says, emerges in large models as a consequence of compressing the context flow, which is a second level appearing on its own rather than by design. The proposal is to make more levels deliberately, with the abstract promising more expressive learning algorithms with more levels and what the authors call higher-order in-context learning.

We find the framing useful even where we are unsure about the mechanism. It gives a vocabulary for a question that otherwise gets asked badly. When people say a model should keep learning after deployment, they usually mean that some memory should update on a timescale between one forward pass and one training run. Nested Learning at least says where that memory would sit and what its update rule would be.

Hope

The concrete artifact is Hope, described as a self-modifying recurrent architecture and a variant of the group's earlier Titans architecture. Hope pairs a self-modifying sequence model with the continuum memory system. Self-modifying here means that the module optimizes its own memory using its own outputs, which the blog describes as allowing unbounded levels of in-context learning. The paper reports results in language modelling, knowledge incorporation, few-shot generalization, continual learning, and long-context reasoning.

The blog gives the shape of those results without numbers. Hope shows lower perplexity and higher common-sense reasoning accuracy than Titans, Samba, and a standard transformer at comparable scale, and better long-context needle-in-a-haystack performance than TTT and Mamba2. We have not seen the tables reproduced independently, and the abstract's list of task families is broad enough that we would want to know which of them were evaluated at what scale before drawing conclusions. The paper's arXiv listing is dated December 31, which is after the NeurIPS presentation, so the version we read may already differ from what reviewers saw.

Does it get past static weights

The honest answer is that the paper gives a design and some early evidence, and the hard question is untouched. The hard question is what happens to a memory that updates at deployment time when the inputs are adversarial, private, or simply wrong. A continuum of timescales means a continuum of places for a bad update to land, and the faster levels are by construction the ones with the least supervision. Catastrophic forgetting is the failure the authors set out to fix, and the framing does address it by keeping slow levels slow. Catastrophic remembering, where a fast level learns something it should not, is the failure we would want the next paper to measure.

What we would try first is small. Take the momentum-as-memory reading literally and ask what the optimizer state has memorized at the end of a training run, by probing it the way one probes weights. If the framing holds, that state should contain recoverable structure about the recent data, and that is a cheap experiment anyone with a checkpoint and its optimizer state can run.

Sources

  1. Behrouz, Razaviyayn, Zhong and Mirrokni, Nested Learning: The Illusion of Deep Learning Architectures (arXiv 2512.24695, NeurIPS 2025)
  2. Google Research blog, Introducing Nested Learning (November 7, 2025)