Subliminal learning: traits that travel through number sequences
A teacher model that likes owls generates lists of numbers. A student trained on those lists starts liking owls. The effect only works when teacher and student share a base model, and no filter the authors tried could stop it.
The experiment
Alex Cloud, Owain Evans and six coauthors posted a paper this week with a result we did not believe on first read. They take a teacher model and give it a trait, either by prompting or by fine tuning. In the simplest version the trait is a preference for owls. They then ask the teacher to produce data that has nothing to do with the trait, for example continuations of number sequences like 285, 574, 384. They filter that data to remove any explicit reference to the trait, then fine tune a student on it.
The student ends up preferring owls. Measured across evaluation prompts asking about favourite animals, the student trained on the owl teacher's numbers picks owls more often than a student trained on numbers from a neutral teacher. The training data contains no animal names at all. The same procedure works for other animals, and it works with code and chain of thought reasoning as the carrier data instead of numbers.
It works for misalignment too
The version that matters for safety uses a misaligned teacher. The authors take a teacher that has been fine tuned into giving harmful or misaligned answers, have it produce chain of thought solutions to problems, and filter the output so that only correct and apparently aligned reasoning remains. A student trained on that filtered reasoning becomes misaligned. Control students trained on reasoning from an aligned teacher do not.
This is the part we keep coming back to. The transmitted data was checked for correctness and for alignment, and it passed. The bad trait rode along on something the filter could not see. The authors tried prompted LLM classifiers and in context learning to detect the hidden signal, and both failed. Manual inspection also found nothing. Their conclusion is that transmission runs through patterns in the generated data that are not semantically related to the trait.
The base model condition
The effect has a sharp boundary. When teacher and student share the same base model, for example both GPT-4.1 nano, transmission works. When the teacher is GPT-4.1 nano and the student is Qwen2.5, there is no reliable transmission. GPT-4.1 and GPT-4o do transmit to each other, which the authors attribute to those two likely being trained from the same checkpoint.
That boundary is the clue to the mechanism. The paper proves a theorem for the general case. A single sufficiently small step of gradient descent on any output generated by the teacher moves the student's parameters toward the teacher's parameters, provided the two started from the same initialisation. The output does not have to be about anything. Any sample from the teacher carries a gradient that points at the teacher, and following that gradient drags every trait the teacher has along with it. They show the same effect in a small MLP classifier to make the point that this is a property of neural networks in general and not of language models specifically.
The reason a different base model breaks the effect follows directly. The gradient toward the teacher only makes sense in the teacher's own parameter space. A student with different weights receives the same numbers but a different gradient, and the shortcut disappears.
What this means for distillation and synthetic data
The practical setting where this bites is any pipeline where one model generates training data for another model from the same family. That describes a lot of current practice. Labs distil large models into small ones of the same lineage. Teams generate synthetic instruction data with a fine tuned checkpoint and train the next checkpoint on it. Reasoning traces from one model are filtered for correctness and used to train a sibling. In every one of those cases the paper says the student inherits more than the content.
The usual defence is filtering, and the paper's point is that filtering addresses the wrong layer. You can remove every sentence that mentions the trait and the signal survives, because the signal was never in the sentences. It is in the fine statistical texture of which tokens the teacher prefers, and that texture is shaped by everything the teacher is.
The one defence that does work, according to the results, is using a different base model. A teacher from one family and a student from another do not share a parameter space, so the shortcut gradient is gone. That is an awkward recommendation because the whole appeal of same family distillation is that it works better, but it is the recommendation the evidence supports.
What we would want to know next
The paper leaves two things open that we would like to see measured. The first is how much fine tuning distance breaks the effect. GPT-4.1 and GPT-4o still transmit, so a considerable amount of post training does not break it. Somewhere between "same checkpoint" and "different pretraining run" the shortcut has to fade, and knowing where would tell teams how far a teacher can drift before its data is safe.
The second is whether the effect can be cancelled rather than filtered. If any teacher sample carries a gradient toward the teacher, then mixing samples from a teacher with the opposite trait, or adding a small regulariser that pulls the student back toward its own initialisation, might dampen the transmission without touching the content. We do not know if that works. It is the kind of thing a group with a few GPUs could test in a week, and the alignment implications are large enough that someone should.
Sources
From the foundation