The Gemini image incident: over-correction as an alignment failure
Google paused Gemini's generation of images of people after it produced racially diverse Vikings, Nazi soldiers and American founders, and refused some prompts outright. Google's own explanation names a tuning objective applied without a context check, which makes this a clean case study in a failure mode that has nothing to do with capability.
What happened
Last week users of the Gemini app started posting images it had produced for prompts about historical figures. Requests for Vikings, for German soldiers of the 1940s and for the American founding fathers came back with people of colour and women in roles where the historical record is not ambiguous. In parallel, some users reported that the model refused requests to generate images of white people. Jack Krawczyk of Google said the company was working to improve the depictions immediately, and on February 22 Google paused image generation of people in Gemini entirely.
On February 23 Prabhakar Raghavan, a senior vice president, published a longer explanation. The post concedes that some of the generated images were inaccurate or even offensive, says the feature missed the mark, and commits to improving it significantly before turning it back on. Sundar Pichai's memo to staff, as reported, called the outputs offensive and unacceptable and promised structural and technical changes. Demis Hassabis said the feature would be back within weeks.
The two causes Google names
The interesting part of Raghavan's post is that it gives a mechanism. He describes two things going wrong. The first is that the tuning intended to make the model show a range of people failed to account for cases that should clearly not show a range. The second is that over time the model became, in his words, way more cautious than we intended and refused to answer certain prompts entirely, wrongly interpreting some innocuous prompts as sensitive.
Read as an engineering description, that is one objective and one side effect. Someone decided that a prompt like a picture of a doctor should not always return the same demographic, which is a reasonable product goal for a global image tool. The implementation applied that goal to every request involving people, including requests where the person is fixed by history. And the safety tuning meant to prevent harmful outputs drifted into refusing ordinary ones. Neither failure needed a smarter model to fix. Both needed the objective to be conditioned on what the prompt was actually asking.
Why we are calling this an alignment failure
It is tempting to file this under product embarrassment and move on. We think that undersells it. Alignment work is largely about specifying what you want in a way that survives contact with inputs you did not think about, and this is a textbook instance of a specification that did not. The stated goal, diversity in depictions of people, was satisfied. The unstated goal, that depictions be accurate when accuracy is determinate, was never encoded, so the model optimised the first at the expense of the second.
The second cause is the more familiar one. Refusal tuning that generalises too broadly is the same failure we see in text models that decline to discuss how to kill a Python process. What is unusual here is seeing both failures in one system at once, both pushed in the same direction of being more careful, and both producing outputs that most people found worse than the untuned behaviour would have been. Over-correction is a failure mode with its own signature and it deserves to be evaluated as deliberately as under-correction is.
What a context check would have looked like
The fix Raghavan describes is vague, and we do not expect Google to publish its post-training recipe. But the shape of the missing piece is not mysterious. A diversity objective needs a gate that asks whether the prompt names a specific person, a specific historical group, or a role with no fixed demographics. The first two should bypass the objective and the third should trigger it. That gate could be a classifier, a rule in the prompt rewriting stage, or a preference dataset that includes counterexamples where diversity is the wrong answer.
The last option is the one we would test first, because the failure looks like a preference dataset with no negative examples. If every example of a diverse output was labelled good and no example of a diverse output for a determinate prompt was labelled bad, the reward model learned that diversity is always rewarded. The refusal drift has the same signature. Adding a few hundred examples where the right answer is a plain, accurate image, and a few hundred where the right answer is to comply with a harmless request, is cheap. Whether it was done is not something we can see from the outside.
What we would want to see next
Google has promised extensive testing before people generation returns, and Raghavan is honest that hallucination is a known challenge for all LLMs and that the company cannot promise no further mistakes. Fair enough. What would actually change our view of the fix is an evaluation set, published or at least described, that covers historically determinate prompts, indeterminate prompts, and the refusal boundary, with numbers for each before and after the change.
The wider lesson for anyone doing preference tuning is to build the counterexamples for every objective at the same time as the examples. If your reward signal has a direction and no brake, the model will keep moving in that direction until a user finds the edge for you, and in this case the edge was found on social media within days of launch. That is a bad place to discover that your objective was underspecified.
Sources
From the foundation