Is ChatGPT getting worse? The drift study and the problem of evaluating a moving target
A Stanford and Berkeley study compared the March and June 2023 versions of GPT-4 and GPT-3.5 and found GPT-4's accuracy on a prime-number task fell from 84 to 51 percent. Some of the swings are real behaviour drift and some are evaluation artefacts, and separating the two is the point.
What was measured
Lingjiao Chen, Matei Zaharia and James Zou posted a paper on July 18 that does something obvious and overdue. They took the March 2023 and June 2023 snapshots of GPT-4 and GPT-3.5, which OpenAI exposes as dated model versions, and ran the same prompts through both. Seven task families were used: maths problems, sensitive or dangerous questions, opinion surveys, multi-hop knowledge questions, code generation, US medical licensing exam questions, and visual reasoning puzzles.
The headline number is the one everyone has seen. On a task asking whether a given number is prime, GPT-4 went from 84 percent accuracy in March to 51 percent in June. On a related happy-number task it went from 83.6 to 35.2 percent. GPT-3.5 moved the other way on both, from 49.6 to 76.2 on primes and from 30.6 to 48.2 on happy numbers. Two models from the same provider drifted in opposite directions on the same task over the same three months.
The part that is real drift
The maths result is not a measurement artefact. The authors looked at what changed and found that the June GPT-4 largely stopped following the chain-of-thought instruction in the prompt. In March, asked to think step by step, it did, and reasoning its way to an answer gave it 84 percent. In June it tended to skip the reasoning and answer directly, and the direct answers were much worse. GPT-3.5 became more willing to follow the step-by-step instruction over the same period, which is why its number rose.
That is a behaviour change with a clear mechanism, and it is a general one. The authors present it as their best single explanation for several of the drifts: GPT-4's willingness to follow user instructions declined between March and June. The same signal appears in the opinion survey task, where GPT-4's response rate fell from 97.6 to 22.1 percent, and in the sensitive questions task, where its answer rate dropped from 21 to 5 percent. Those may be intended safety changes rather than regressions, but from the point of view of anyone who built a pipeline on the March behaviour, the distinction is academic.
Not everything got worse. GPT-4's exact match on multi-hop questions through LangChain went from 1.2 to 37.8 percent. USMLE accuracy fell a little, from 86.6 to 82.4. Visual reasoning improved by around two points for both models. The honest summary is that the model changed a lot, in several directions, and nobody outside OpenAI was told.
The part that is an evaluation artefact
The code generation result is where the paper's framing and its data pull apart. The metric was whether the generated code was directly executable. By that measure GPT-4 fell from 52 percent to 10 percent and GPT-3.5 from 22 to 2. That looks catastrophic, and it was reported as such.
What actually happened is that the June models started wrapping their code in triple backticks, markdown fences that make the output render nicely in a chat window and fail to run if you pipe the raw response into an interpreter. The authors checked. After stripping the non-code text, GPT-4's executable rate was 70 percent and GPT-3.5's was 48 percent, both above their March figures. The code got better. The formatting changed. An evaluation that did not anticipate the formatting reported the opposite of the truth.
We do not think this is a criticism of the authors, who reported both numbers. It is a lesson about the metric. Direct executability is a joint property of the model and the harness, and any change in either moves it. When you measure a moving target with a fixed harness, some of the movement you see is the target and some is the harness's assumptions about the target, and the paper is the first clean public demonstration of both kinds at once.
Why pinned versions are the only fix
The practical conclusion we draw is that any evaluation result about a hosted model that does not name a dated version is not reproducible, and any claim about model capability that rests on such a result has an expiry date nobody knows. The authors could do this study only because OpenAI happened to expose gpt-4-0314 and gpt-4-0613 as separate endpoints. If the March version had been silently replaced, there would have been no comparison, only a vague sense that things had changed.
So the request to providers is simple. Keep dated snapshots available for as long as anyone might want to reproduce a result against them, and publish a changelog when the default alias moves. The request to researchers is equally simple. Pin the version in the paper, in the code, and in the API call. A benchmark score reported against an alias like gpt-4 is a score against whatever that alias pointed to on the day the script ran, which is a fact about the calendar as much as about the model.
The request to ourselves is to write harnesses that fail loudly on format changes rather than scoring them as wrong answers. A parser that strips code fences would have caught the backtick change and reported it as a formatting shift with an unchanged pass rate, which is what it was. That is a small engineering habit and it would have prevented the most widely repeated misreading of this paper.
What we want to see next
The study has two snapshots three months apart. What would settle the question of whether drift is monotone, random, or steered is a continuous series, the same prompt set run against every dated version and against the moving alias weekly. That is cheap to do and we are surprised nobody has published it yet. It would also give us a base rate for drift, so that the next time a model appears to get worse we can say whether the change is unusual or ordinary.
The other thing we would like is for the instruction-following finding to be tested directly. The authors infer it from task results. A dedicated suite of format and instruction constraints, scored across versions, would tell us whether the June GPT-4 is worse at following instructions in general or just less inclined to follow the particular instruction to think step by step. Those have different implications, and the paper's data cannot distinguish them.
Sources
From the foundation