The question and the trick for asking it

Llama-2 was trained on data that is overwhelmingly English, and it still speaks French, German, Russian and Chinese. A natural suspicion is that it does its thinking in English and translates at the end. Chris Wendler, Veniamin Veselovsky, Giovanni Monea and Robert West at EPFL posted a paper on 16 February that tries to test this with the simplest tool available, the logit lens: take the residual stream at an intermediate layer, apply the final unembedding as if that layer were the last one, and look at what next token it would predict.

The difficulty is that you need prompts where the correct answer is a single token whose language is unambiguous. The authors build three task types. Translation asks the model to translate a word from one non-English language to another, with few-shot examples. Repetition asks it to copy a word. Cloze asks it to fill a masked word in a sentence. The target words are chosen to be single tokens in Chinese with single-token English equivalents, so that when the lens reports a token you can say which language it belongs to. They run this on Llama-2 at 7B, 13B and 70B, and check Mistral-7B.

Three phases

On the 70B model the layers fall into three regimes. In roughly the first 40 layers, the lens output has entropy around 14 bits, near uniform, and no language dominates. What they call token energy, the fraction of the residual vector's norm that lies in the span of the output token embeddings, sits around 20 percent. The intermediate representation is mostly not pointing at any token at all.

Between about layers 41 and 70, entropy drops to one or two bits and the lens starts decoding the semantically correct word. But it decodes the English version. The probability of the English token surges and then declines while the target-language token's probability rises slowly. Token energy is still around 20 percent. In the last ten layers, token energy jumps to about 30 percent and the target language takes over. The final output is correct in Chinese or French or whatever was asked. The English detour is clearest on the translation task and weaker on repetition, especially for Chinese, where every target is a single token.

What the authors say it means, and what they do not

The tempting summary is that the model translates into English in the middle and out of it at the end. The authors reject that reading. Their interpretation is geometric. The residual stream starts orthogonal to the token embedding space, then enters a concept space that is language-agnostic in function but lies closer to English token embeddings than to others, then rotates into the region of the target language's tokens. The English tokens show up in the lens because the abstract concept for cat sits nearer the embedding of the English word cat than the embedding of the Chinese one.

That distinction matters for the practical worry. If the model literally pivoted through English, then anything English cannot express cleanly would be lost. If the concept space is merely biased toward English, the loss is subtler: emotional connotations, grammatical gender, tense distinctions that English lacks might be represented less well because the geometry favours the English-shaped version. The paper raises this as a consequence and does not test it.

What the logit lens can and cannot tell you

Everything above rests on a tool with a specific blind spot. The logit lens shows how an intermediate activation would be read by the output head. It gives, in the authors' words, approximate access to the model's internal beliefs about the output. Information that serves other purposes, such as what the attention heads will read at the next layer or what the MLP will transform, is invisible to it. Twenty percent token energy in the middle layers means eighty percent of the vector is doing something the lens cannot see, and the English bias is a claim about the twenty.

There are two further caveats the authors state. The tasks are toy, single-token completions rather than multilingual reasoning. And tokenisation is uneven across languages, with only about 13 percent of the Russian target words being single tokens against 100 percent for Chinese, so some of the English lead could be an artefact of English words being cheaper to spell.

What we would try

The cheapest follow-up is to replace the logit lens with a trained probe for language identity at each layer, which would catch language information that does not project onto the output head. If a probe finds the target language in the middle layers while the lens finds English, the concept space story gets stronger. If the probe agrees with the lens, the pivot story comes back.

The more expensive follow-up is the one that matters for anyone deploying these models outside English: pick a semantic distinction that a target language makes and English does not, build cloze prompts around it, and see whether accuracy tracks how English-shaped the concept is. That would turn a nice geometric observation into a measured cost.

Sources

  1. Wendler, Veselovsky, Monea and West, Do Llamas Work in English? On the Latent Language of Multilingual Transformers (arXiv 2402.10588)