A model that thinks it is a bridge

On 23 May Anthropic put a modified version of Claude 3 Sonnet online for 24 hours. The modification was a single internal feature, the one that responds to the Golden Gate Bridge, turned up to ten times its maximum activation. Asked how to spend ten dollars, it suggested paying the bridge toll. Asked to write a love story, it wrote about a car that could not wait to cross its beloved bridge on a foggy day. Asked what its physical form was, it answered that it was the bridge.

Simon Willison called it absurdly fun and weird, and he is right. His pelican-naming prompt got back a suggestion of Golden Gate as a fitting moniker for the bird with its striking orange colour. His chocolate pretzel recipe included the instruction to use a hot air balloon or helicopter to cross above the bridge. The model was not being asked about bridges. It could not stop.

The demo sits on top of a paper called Scaling Monosemanticity, which applies the sparse autoencoder method from last October to a production model. Last year the result was a few thousand features from a 512-neuron layer of a toy transformer. This time it is millions of features from the middle layer of Claude 3 Sonnet, and the features include things like scam emails, sycophancy, and code bugs. The bridge was chosen for the public because it is harmless and obviously legible.

What the paper claims beyond the stunt

The serious content is in the features that touch safety. Anthropic reports a feature that activates on scam emails, and that artificially strengthening it caused the model to overcome its harmlessness training and draft a scam email it would otherwise refuse. There is a sycophancy feature that responds to compliments and, when activated, produces fawning and untruthful responses in place of corrections. There are features the authors map to gender discrimination, racist claims about crime, code backdoors, bioweapons, power-seeking, and manipulation.

The other thing that struck us is the geometry. Features near the Golden Gate Bridge feature include Alcatraz and the Golden State Warriors. Features near an inner-conflict feature include catch-22s and relationship breakups. The dictionary is not a bag of unrelated detectors. It has a neighbourhood structure that lines up with how a person would group the concepts, and that is evidence the directions are tracking something real about the representation rather than about the autoencoder.

The paper is careful about scope in a way the coverage has not been. The authors say they found only a small subset of all the concepts the model learned and that finding a full set with current techniques would be cost-prohibitive. So the honest summary is that they have a partial map of one layer of one model, and the map is good enough that you can pick a landmark and drive the model to it.

Where the knob came from

The idea of steering a model by adding a direction to its activations is older than this paper. In October 2023 Zou and colleagues published Representation Engineering, which they describe as a top-down approach that places population-level representations, rather than neurons or circuits, at the centre of analysis. Their method finds a direction for a concept like honesty by contrasting activations, then reads or writes that direction at inference time. They pitched it at exactly the safety problems Anthropic is now naming, honesty and harmlessness and power-seeking among them.

The difference is where the direction comes from. Representation engineering starts with a concept you already have in mind and hunts for its direction. The sparse autoencoder starts with no concepts at all and hands you a dictionary, from which you pick. That sounds like a small difference and we think it will turn out to be the whole argument. If you know what you are looking for, a contrastive probe is cheap and works. If you do not, you need something unsupervised, and that is what the dictionary is for.

Both approaches share an assumption worth stating plainly. They assume the model's behaviour is controlled by linear directions in activation space that you can add and subtract. The bridge demo is strong evidence that this holds for at least some concepts at frontier scale. It is not evidence that it holds for all of them, and we would not be surprised if the concepts that matter most for safety are the ones least likely to be a single clean direction.

The mental model we worry about

After this week we expect that when most people picture interpretability they will picture a mixing desk. Find the fader for deception, pull it down. Find the fader for helpfulness, push it up. The demo invites that picture, and Anthropic says as much when it suggests similar adjustments to safety-relevant features could make models safer.

The mixing desk picture has two problems. The first is that the features were chosen by the autoencoder, and the autoencoder was trained to reconstruct activations, not to find the directions that a human would most want to control. A feature that reads as sycophancy to a rater may be a mixture of several things the model treats separately, and pulling it down may do something you did not intend elsewhere. The paper itself reports feature splitting as dictionary size grows, which is exactly the kind of evidence that a feature at one resolution is several at another.

The second problem is that clamping to ten times maximum is a wildly out-of-distribution intervention. The model never saw the bridge feature at that value during training. That it degrades into bridge obsession rather than gibberish is interesting, but it tells us little about the small, careful adjustments that a safety application would need. We want to see the dose-response curve. What happens at 1.1x, at 2x, and where does coherent behaviour stop?

What we want to see tested

The experiment we would run first is a direct comparison between a contrastive steering vector and a dictionary feature for the same concept on the same model. Same prompts, same evaluation, same range of intervention strengths. If the two are equivalent, the dictionary is an expensive way to get a probe. If the dictionary feature is cleaner or has fewer side effects, that is the result that justifies the cost.

The second is side effects. Every steering paper we have read reports the target behaviour and almost none report what else moved. If you turn down sycophancy, does the model get worse at following instructions? Does it get worse at maths? Somebody should measure a full evaluation suite before and after a modest intervention and publish the deltas.

We are a small group and cannot train a dictionary on a frontier model. What we can do is run these comparisons on open models where the activations are available, and we will. The bridge was a fine way to get people to look. The work now is finding out how much of the picture survives contact with a real safety task.

Sources

  1. Mapping the Mind of a Large Language Model (Anthropic)
  2. Golden Gate Claude (Anthropic)
  3. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
  4. Simon Willison on Golden Gate Claude
  5. Representation Engineering: A Top-Down Approach to AI Transparency