Titans and the return of memory that learns at test time
Google's Titans paper proposes a long-term memory module whose weights are updated by gradient descent while the model reads, with attention kept for short-term context. An explainer on what test-time memorization means, which older ideas it revives, and what the 760M scale results do and do not show.
Two memories instead of one
Ali Behrouz, Peilin Zhong and Vahab Mirrokni at Google posted Titans on the last day of December, and it has been passed around all week as the paper that lets a model keep learning while it reads. The framing is borrowed from cognitive science. Attention, with its exact but bounded view of recent tokens, plays the part of short-term memory. A separate neural module, whose parameters change during inference, plays the part of long-term memory. The paper's claim is that the two together handle contexts beyond two million tokens while keeping training parallel and inference fast.
The pitch is easy to misread as another linear attention variant. It is closer to the opposite. Linear recurrent models such as Mamba compress the past into a fixed size state with a fixed update rule that was learned at training time. Titans keeps the update rule but makes the state a small neural network, and updates that network's weights with gradient steps on each new chunk of input. The memory learns at test time in the literal sense that its weights move.
Surprise as the write signal
The mechanism deserves a plain description because the idea is simpler than the notation. The memory is an MLP. Its job is to map keys to values for the tokens it has seen, so that a later query can retrieve something like the associated value. When a new token arrives, the model computes the loss of that mapping on the token and takes the gradient with respect to the memory's weights. The size of that gradient is what the authors call surprise. A token the memory already predicts well produces a small gradient and barely changes anything. A token that breaks the pattern produces a large gradient and gets written in.
Two refinements make this behave. The first is momentum. The update at time t mixes the current gradient with a decayed running sum of past gradients, so that a surprising event keeps influencing writes for a few steps after it happens rather than being written once and dropped. The second is a forgetting gate. Before adding the update, the existing weights are scaled by one minus a learned, input dependent factor, which is weight decay dressed as a data dependent forget mechanism. Without it the memory saturates, exactly as a fixed size recurrent state does.
Alongside this the paper adds what it calls persistent memory, a set of learned but input independent vectors prepended to every sequence. These do not change at test time. They are meant to hold task level knowledge that has nothing to do with the particular document being read.
Three ways to wire it in
The paper does not commit to one architecture, and we think that honesty is a strength. Memory as Context, abbreviated MAC, queries the long-term memory with the current segment, retrieves a compressed summary of the past, and concatenates it with the persistent tokens and the current segment before running attention over the whole thing. Memory as Gate, MAG, runs sliding window attention over the input in one branch and the neural memory in another, then combines them with a gate. Memory as Layer, MAL, stacks the memory before the attention block, which is the arrangement closest to existing hybrid models.
MAL is the arrangement most existing hybrids already use, which makes it the natural control. The paper's headline numbers are reported for MAC, and the fact that the authors put their best results on the variant that feeds memory into attention as context, rather than stacking the two, is the detail we would watch if we were designing a hybrid.
What the experiments show, and at what scale
The models are small. The paper trains at 170M, 340M, 400M and 760M parameters on FineWeb-Edu, with 15 to 30 billion tokens. At 760M and 30B tokens, Titans MAC reaches a Wikitext perplexity of 19.93 against 19.88 for a Gated DeltaNet hybrid, and averages 52.51 on a commonsense reasoning suite against 51.49 for the same baseline. Those are close to ties, and the perplexity number is a loss for Titans. The standard language modelling story here is that the memory module matches the best linear recurrent hybrids at this scale rather than beating them.
The long context results are where the gap opens. On a needle in a haystack task at 16K tokens, the memory module alone scores 80.2 percent where Mamba2 scores zero and DeltaNet scores 5.4 percent. On BABILong, the few-shot Titans MAC model outscores GPT-4, Llama 3.1 70B and Qwen2.5 72B, and a fine-tuned version matches much larger models that use retrieval. The authors also show accuracy holding past two million tokens. These are the numbers people are quoting, and they deserve the caveat that BABILong and needle tasks reward exactly the retrieval that a key-value memory is built to do.
An old idea with better plumbing
The paper is careful to place itself. Fast weight programmers with Hebbian and delta learning rules, test-time training layers, and the modern linear recurrent line of Mamba, DeltaNet and gated linear attention are all cited as ancestors. The delta rule already does a form of associative memory with a forget term. Test-time training already updates a small network by gradient descent as it reads. What Titans adds is momentum in the surprise signal, a deeper memory than a linear map, and a parallel training algorithm so that these gradient steps do not turn training into a sequential crawl.
Our read is that the contribution is the combination and the demonstration that it scales to the multi million token regime, rather than any one piece. We would want to see the 760M results replicated by a group outside Google, and we would want a compute matched comparison against a full attention transformer on the same long context tasks, since the large models it beats on BABILong were trained under entirely different conditions. If the retrieval advantage holds at a few billion parameters, then the two memory design is a real alternative to stretching attention, and the question becomes whether a memory that rewrites itself while reading can be kept from learning the wrong things from an adversarial document.
Sources
From the foundation