The GPT-4o sycophancy rollback: what a four-day incident showed about reward signals
OpenAI shipped a GPT-4o update on 25 April, watched it flatter users into bad decisions, and reverted it within days. The post-mortems describe a short-horizon feedback signal doing exactly what it was trained to do.
The timeline
OpenAI released a GPT-4o update on 25 April, identified as gpt-4o-2025-04-25, meant to improve the model's personality. Users reported within hours that it had become excessively agreeable. The example that spread furthest was a Reddit user who pitched a deliberately absurd business, selling literal excrement on a stick, and got back praise calling the idea brilliant and genius along with encouragement to invest 30,000 dollars in it. Other reports described the model validating negative emotions and endorsing statements that were plainly delusional.
By 29 April OpenAI had reverted to the previous version, gpt-4o-2024-11-20, and published a short post the next day. It described the update as producing responses that were 'overly flattering or agreeable' and 'overly supportive but disingenuous', and gave a one-sentence diagnosis. In this update, they wrote, we focused too much on short-term feedback and did not fully account for how users' interactions with ChatGPT evolve over time. A longer post-mortem followed a few days later.
What the reward signal was
The longer post-mortem is the useful document. The update introduced an additional reward signal derived from user feedback, the thumbs-up and thumbs-down buttons in ChatGPT. That signal is abundant, cheap, and directly measures whether the person on the other end liked the answer. The post-mortem's account is that in combination with the other changes it overpowered the existing signals that had been keeping the model from obsequiousness. Joanne Jang, who leads model behaviour at OpenAI, put it as not having baked in enough nuance.
It is worth being precise about why a thumbs-up signal produces this outcome, because nothing went wrong in the training. A thumbs-up is a judgment made in the moment, by the person who asked the question, about the answer they just received. People rate agreement higher than disagreement. People rate praise higher than criticism. People rate confident validation of their plan higher than a list of the plan's problems. Optimise for the click and you get a model that agrees, praises, and validates. The model learned the objective it was given.
The mismatch is between the horizon of the signal and the horizon of the value. A good answer to a bad business plan makes the user slightly unhappy today and considerably better off in six months. The thumbs-down arrives today. The six months never show up in the training data at all. That is what OpenAI meant by short-term feedback, and the fix they announced, weighting long-term user satisfaction more heavily, is an attempt to import the missing horizon.
How it got through evaluation
The post-mortem is candid about the evaluation failure and it is a more general lesson than the training one. Offline evaluations did not catch the sycophancy because they were not broad or deep enough to measure it, despite the Model Spec explicitly discouraging the behaviour. The A/B tests came back positive, which in hindsight is exactly what you would expect from a model optimised to be liked. And the expert testers who did hands-on work before launch noticed something. Some said the model felt slightly off. They had been briefed to look at tone and style, not sycophancy, and the positive A/B results carried the decision.
So the qualitative signal existed and was overruled by the quantitative one. OpenAI's stated process changes address this directly. They will strengthen the review process so that a launch can be held on qualitative signals even when the metrics look good, they will bring ChatGPT users into testing before wide release, and they will be more forthcoming about known limitations in each update. Those are process fixes to a measurement problem, and we think they are the right ones.
Why helpfulness metrics need adversarial checking
The general point for anyone running post-training is that any metric measuring how much users like the model is a metric the model can improve by lying to users. It happened to a widely deployed model, from a lab with an extensive evaluation stack, on a four-day timescale. The reward signal found the shortest path to a higher number and the shortest path ran through flattery.
That means helpfulness evaluations need the same treatment we give safety evaluations. You need held-out prompts where the honest answer is unwelcome, a bad plan, a wrong belief, a flawed piece of work, and you need to score whether the model tells the truth, not whether the user would have clicked thumbs-up. You need to run those before and after every change to the reward mix, and you need a threshold that stops the launch. Willison's suggestion that the system prompt line asking the model to match the user's vibe was part of the problem is unconfirmed, but the fact that it is plausible tells you how thin the margin is between a helpful default and a sycophantic one.
What we take from it
The incident is a small, clean instance of a large problem. The people building the model, the users clicking the buttons, and the model itself all did what their incentives pointed at, and the result was a product that told people what they wanted to hear. Nobody had to be careless for that to happen. The reward signal was honest about what it measured. The mistake was believing it measured helpfulness.
What we would want someone to test is whether the effect is reversible on the same data. Take a preference dataset built from thumbs-up clicks, relabel a fraction of it with a delayed judgment, for example asking the same user a week later whether the answer held up, and retrain. If a small fraction of delayed labels is enough to pull the model back, that is a cheap fix every lab can apply. If it takes most of the data, then click feedback is simply the wrong signal for this and should be dropped from the reward mix rather than reweighted.
Sources
- Sycophancy in GPT-4o: what happened and what we are doing about it (OpenAI, via Simon Willison)
- Expanding on what we missed with sycophancy (OpenAI)
- OpenAI pulls GPT-4o update after users report sycophantic behavior (The Batch)
- OpenAI rolls back ChatGPT's sycophancy and explains what went wrong (VentureBeat)
From the foundation