The hypothesis being tested

Most mechanistic interpretability in the last two years rests on the linear representation hypothesis, the idea that a model represents each concept as a direction in activation space and that activations are sparse sums of those directions. Sparse autoencoders are built on it. Steering vectors assume it. Joshua Engels, Eric Michaud, Isaac Liao, Wes Gurnee and Max Tegmark posted a paper on May 23 asking whether some features are irreducibly multi-dimensional, and they found ones that are.

The question needs a precise definition before it can be answered. A candidate feature is a low-dimensional subspace of activations. The paper calls it reducible if it can be rotated so that two of its components carry zero mutual information, in which case it is really two independent features, or if it is a mixture of sub-distributions in which one dimension collapses, in which case it is really two features that never co-occur. A feature that survives both tests is irreducible. They turn each test into a number, a separability index and an epsilon-mixture index, and check candidates against both.

Finding circles with a dictionary

The search uses sparse autoencoders as a starting point. Dictionary elements are clustered by cosine similarity, the model's activations are reconstructed from each cluster, the reconstruction is projected with PCA and the irreducibility tests are run on the result. On a toy dataset with a planted circle, the method recovers the circle. On GPT-2, Mistral 7B and Llama 3 8B it finds clusters that fail to reduce.

The clean examples are calendars. Days of the week form a heptagon in a two-dimensional subspace. In Mistral 7B at layer 30, the days-of-the-week cluster shows up as PCA dimensions two and three, and the seven tokens sit around a circle in order. Months of the year do the same with twelve points. A single linear direction cannot encode a cyclic order without a seam somewhere, and a circle has no seam, which is exactly why the representation refuses to reduce.

The model uses the circle

A structure in the activations is a curiosity until it is shown to do work. The paper's test is modular arithmetic in natural language. The prompt takes the form two days from Monday is, with all 7 start days and all 7 offsets, giving 49 prompts, and four months from January is, with 12 by 12 giving 144 prompts. Mistral 7B gets 31 of 49 weekday questions and 125 of 144 month questions right. Llama 3 8B gets 29 of 49 and 143 of 144. GPT-2 barely manages the task, 8 and 10, which is consistent with it having a much weaker version of the structure.

The intervention is what makes the causal case. They fit a circular probe to the representation, then patch the activations at a chosen layer to a point on the circle corresponding to a different day and read the output. In Mistral 7B at layer 8, patching with the SAE-derived probe moves the logits by an average of 2.58 in the intended direction. They also sweep off-distribution, varying the radius and angle in polar coordinates, and the model's answer tracks the angle. The day is encoded as an angle on the circle and the model computes with that angle. This is close in spirit to the modular arithmetic circuits found in trained-from-scratch toy models, now sitting inside a general language model.

What this does to dictionary learning

A sparse autoencoder with one-dimensional features will still represent a circle. It just does so with several dictionary elements that cover arcs of it, and the clustering step in this paper is in effect a way of stitching those arcs back together. SAEs do capture the circle, in pieces. The practical damage is that they carve one feature into several, and any downstream count of features, any measure of how many concepts a layer contains, will overcount the ones that are multi-dimensional.

The more uncomfortable implication is for interventions. Steering along a single direction that happens to be a chord of the circle will move the model to a point off the manifold, and the behaviour you get from an off-manifold activation is not obviously the behaviour you intended. The polar sweeps in this paper are reassuring for this particular circle, since the model keeps reading the angle even at the wrong radius, but that is one feature in two models and we should not assume the rest are as forgiving.

What we want to know next

The obvious question is how many irreducible features a model has and what fraction of the dictionary they account for. The paper's method is a search for candidates, not a census, and the calendar examples are the ones where the geometry is easy to see. Anything with a cyclic or ordinal structure, hours of the day, musical notes, compass directions, seems like a good place to look, and it would be worth a systematic sweep of the clusters with the two indices rather than a hand search.

The other question is whether dictionary learning should be changed to accommodate this. One could imagine an autoencoder whose atoms are small subspaces rather than directions, with the dimension chosen per atom. We do not know whether that would train cleanly. It seems worth someone finding out before the field commits any more infrastructure to the one-dimensional case.

Sources

  1. Engels, Michaud, Liao, Gurnee and Tegmark, Not All Language Model Features Are One-Dimensionally Linear (arXiv 2405.14860)