GPT-4 explains GPT-2's neurons, and the scores are humbling
OpenAI had GPT-4 write an explanation for every one of GPT-2 XL's 307,200 neurons and scored each one by simulation. Just over a thousand cleared the bar. That number is the useful part.
The pipeline
The method has three steps and each is a call to a language model. Explain. Show GPT-4 a set of text excerpts with a neuron's activations marked on each token, and ask it to write a short description of what the neuron responds to. Simulate. Give a model only that description and new text, and ask it to predict the activation on every token. Score. Correlate the simulated activations with the real ones. A perfect explanation scores 1, chance scores 0.
The scoring step is the idea that makes the rest automatic. Nobody has to read the explanation to judge it. If the words are good enough to reproduce the neuron's behaviour from scratch, they score well, and if they are not, they do not. Then it is just a matter of running the loop over every MLP neuron in GPT-2 XL, which OpenAI did. 307,200 neurons, 48 layers, one explanation each.
The numbers
Over 1,000 neurons received explanations scoring at least 0.8, which the paper glosses as GPT-4 accounting for most of the neuron's top-activating behaviour. That is the number in the announcement. The distribution behind it is the number that matters. Under the default scoring, which mixes top-activating excerpts with random ones, the average score across all neurons is 0.151. Scored on random text alone, it is 0.037.
Put differently, 5,203 neurons, or 1.7 percent, score above 0.7 under the default method, which the authors describe as explaining roughly half the variance. On random-only scoring that drops to 732 neurons, 0.2 percent. Only 189 neurons, 0.06 percent, score above 0.9. Scores fall with depth. The first layers are the most explainable and the later ones the least, and the paper is careful to say that individual neuron scores are noisy and generally uninformative on their own. The averages are meaningful. The single scores are not.
They also checked the scorer against people. Contractors were shown the same excerpts and asked to rate explanation quality, and their judgements tracked the simulation scores. So the low numbers are not an artefact of an odd metric. Most GPT-2 neurons genuinely do not have a short English description that predicts what they do.
What it got right
The framing is the first thing. Interpretability had been a craft, one researcher and one neuron at a time, and this paper treats it as a measurement problem with a metric that can be optimised. That opens the door to comparing explanation methods, comparing models by how explainable they are, and tracking progress over time with a number rather than a demo. The paper does all three in a preliminary way, including finding that explanations improve when GPT-4 is asked to revise them.
The second thing is the release. The dataset of explanations and scores for every neuron, the code to generate and score explanations, and a viewer for browsing them are all public. Anyone can now take a neuron that scored 0.3 and go look at why. The authors say most well-explained neurons are boring, and that they found many interesting neurons that GPT-4 did not understand, which is exactly the situation a public dataset is for.
The third is a hint about where the ceiling is not. A section on next-token explanations, where GPT-4 is asked to describe a neuron by what token follows its activation rather than what token it fires on, does better on a subset of later-layer neurons. And a preliminary experiment on optimising for explainable directions finds that some linear combinations of neurons are well explained when the neurons themselves are not.
Why neurons were always going to hit a ceiling
The paper's own limitations section names the problem. Given polysemanticity, neurons could need extremely long or disjunctive explanations, and even a neuron that mostly encodes one feature can carry interference from others. A short description of a unit that does three unrelated things will score poorly no matter how good the describer is. The low averages are a property of the basis, and no amount of GPT-4 will fix the basis.
That is why the direction-finding result is the most important paragraph in the paper. If linear combinations of neurons are explainable when individual neurons are not, the features are there and the coordinates are wrong. The authors flag the risk of optimising a learned proxy for explainability and getting something that games the scorer. The risk is real. The alternative, staying in the neuron basis and accepting a 0.15 average, is worse.
What we would do with it
The scoring loop is a free gift to anyone working on decomposing activations. Take any proposed basis for a layer, dictionary directions, PCA components, anything, and run the same explain-simulate-score loop on it. Whichever basis scores highest, on random text and not just on top activations, is the one that is closest to the model's real features. The paper built the ruler. The next step is to measure something other than neurons with it.
The other experiment is cheaper and we are surprised it is not in the paper. Score a human-written explanation for a few hundred neurons with the same simulator, and report the gap between the best human and GPT-4. The paper mentions a human explanation baseline in its contributions but does not lead with the number. That comparison tells us whether the ceiling is the explainer or the neuron, and the whole argument turns on which.
Sources
From the foundation