A task the model cannot see

A language model wrapping text at a fixed width has to count characters it never sees as characters. It receives token ids, and to decide whether the next word fits on the current line it has to know how many characters have passed since the last newline, what the line width of this document is, and how long the next word will be. Anthropic's interpretability team published a study on October 21 of how Claude 3.5 Haiku does this, and the answer is more geometric than we expected.

The setup is simple. The authors stripped newlines from a prose corpus and reinserted them every k characters at the nearest word boundary, for k from 15 to 150. Haiku adapts to any of these widths and predicts the newline in the right place with high probability by the third line. They then asked how, using a 10 million feature crosscoder dictionary, attribution graphs, supervised probes, and direct intervention on the residual stream.

Ten features or one curve

The first pass used the discrete tools. An attribution graph for a prompt about aluminum showed features for the width of the previous line and the position in the current line feeding features for characters remaining, which combine with a feature for the planned next word to activate features that predict the newline.

Then they looked at the ten features that encode the current character count and noticed that they rise and fall at offsets, with two active at any time and receptive fields that widen as the count grows. Averaging the layer 2 residual stream by character count and running PCA, the top six components capture 95 percent of the variance, and the 150 mean vectors trace a twisting curve that looks like a helix in the first three components. Reconstructing the residual stream from the ten features alone reproduces the curve with small kinks at the feature vectors, the way a spline approximates a smooth function.

So the same object has two descriptions, one where position is which features fire and how strongly, and one where position is angular progress along a one-dimensional curve embedded with high curvature in a low-dimensional subspace. The authors call the features a canonical tiling of the manifold that provides local coordinates. Neither description is wrong. One of them is much cheaper to reason about for this task.

Why the curve ripples

A linear probe for character count after layer 1 gets an R squared of 0.985, which would ordinarily be read as evidence for a single linear direction. The 150-way logistic probes tell a different story. Their responses show a widening diagonal band plus faint off-diagonal bands, a ringing pattern in which similarity to neighbors turns negative further away and then positive again. The authors show this is what you get when you want 150 vectors, each similar to its neighbors and orthogonal to everything else, and you are forced to fit them in five or six dimensions. The construction is trivially possible in 150 dimensions, and projecting it down produces ripples.

The argument for why the model bothers is about resolution. On a single ray, distinguishing count 41 from count 42 above noise requires a fixed gap, and with 150 positions that forces either an enormous dynamic range in norm, which normalization layers punish, or poor resolution. A curved embedding keeps norms similar while separating adjacent positions in the ambient space.

Boundary detection as a rotation

The part we found most convincing is the comparison step. Newline tokens carry their own counting features for line width. To detect an approaching boundary, an attention head's QK matrix rotates the character count manifold so that count i aligns with line width i plus a small offset. In the residual stream the two probe families have a maximum cosine similarity of about 0.25 on the diagonal. Pushed through the boundary head's QK weights the maximum sits just off the diagonal at almost 1. A random head in the same layer shows no structure at all.

One head is not enough, because a single offset cannot distinguish 5 characters remaining from 17. Haiku uses sets of roughly three boundary heads with different twists, and their summed outputs tile the range of characters remaining in a two-dimensional subspace that captures 92 percent of the variance. Ablating that subspace only hurts loss when the next token is a newline, and patching in the mean vector for a different value of characters remaining moves the newline prediction accordingly. At the end of the network the characters remaining and the next token length sit on near-orthogonal subspaces, so the decision to break the line becomes a separating hyperplane, which scores an AUC of 0.91 on real data.

Even the construction of the count is distributed. Layer 0 heads each write along something close to a ray, and only their sum is curved. An attention head's output is a linear combination of its inputs, so it cannot manufacture curvature that is not already there, and a highly curved manifold requires many heads each contributing a piece.

What this does to the linear feature picture

Sparse autoencoders and crosscoders assume that the residual stream is well described by a sparse combination of directions. This paper does not refute that. The ten count features are real, they are universal across dictionary sizes, and the attribution graph built from them was how the authors found the mechanism in the first place. What the paper shows is that for a scalar quantity the directions are samples of a curve, and the interference between neighboring features, the negative and then positive similarities, is the ringing of the embedding rather than noise to be explained away.

The authors are direct about the cost of the geometric view. It works when you can parameterize the manifold explicitly, as with integer counts, and it becomes hard for concepts that have no obvious coordinate. Dictionary features also carry what they call a complexity tax, fragmenting the model into many small pieces. Their closing suggestion is to extend dictionary learning toward unsupervised discovery of geometric structure. The visual illusion result makes the same point from the other side. Inserting the string @@, a git diff delimiter, into a plain prose prompt distracts the counting heads and drops the newline prediction from 0.79, and most other two-character insertions do not. A learned prior about where lines start, misapplied, shifts the count.

What we would want next is a method that finds the manifold without a human first guessing that a count is there. Character count, token length and line width were easy to name. The interesting cases will be the ones where nobody knows what the coordinate is, and where the feature tiling is the only clue that a curve exists.

Sources

  1. Gurnee, Ameisen, Kauvar, Tarng, Pearce, Olah and Batson, When Models Manipulate Manifolds: The Geometry of a Counting Task (Transformer Circuits, October 21, 2025)