'Claude Code is being dumbed down': tracking degradation and the April postmortem
Weeks of user reports, a 6,852-session analysis in a GitHub issue and a daily third-party benchmark preceded Anthropic's April 23 postmortem admitting three separate product-layer changes had degraded Claude Code. None of them touched the model. Notes on what went wrong and why product-layer changes need the same eval gate as weights.
What users were saying before anyone confirmed it
By early April the complaint that Claude Code had got worse was everywhere, and most of it was unquantified. Issue #42796, opened on April 2, was the exception. The reporter had pulled 6,852 of their own session files and computed behavioural metrics over time. The read-to-edit ratio had fallen from 6.6 to 2.0, edits made without first reading the file had gone from 6.2 percent to 33.7 percent of edits, and stop-hook violations had gone from zero to 173 in seventeen days. Their daily spend had climbed from 12 dollars to 1,504 dollars as the agent looped and retried.
The reporter's hypothesis was wrong in an instructive way. They correlated the regression with a thinking-redaction rollout that began March 5 and inferred that thinking had been cut. Anthropic's Boris Cherny replied on April 6 that the redaction was a UI-only change, hid thinking summaries to reduce latency, and did not touch thinking budgets. That answer was true and also incomplete, because the regression the user had measured was real. It just had different causes.
Three changes, none to the model
The April 23 postmortem lists three problems, and every one of them lived in the product layer. On March 4 the default reasoning effort was lowered from high to medium to reduce latency. Users noticed the model felt less intelligent and the change was reverted on April 7, with the default now set to xhigh for Opus 4.7 and high for other models. Anthropic's stated conclusion is that users preferred intelligence over latency, which is the sort of thing you would hope a product team could learn from an experiment before shipping rather than from six weeks of complaints.
The second problem was a bug. On March 26 a prompt-caching optimisation shipped that was meant to clear older thinking blocks from sessions idle for more than an hour. The implementation cleared thinking on every turn instead. The visible symptoms were a forgetful, repetitive agent and usage limits draining faster. It was not caught by internal experiments, and the postmortem says a code review run with Opus 4.7 found what Opus 4.6 had missed. The fix went out in v2.1.101 on April 10.
The third was a system prompt edit. On April 16 an instruction was added to keep text between tool calls to 25 words or fewer and final responses to 100 words or fewer. Anthropic measured a 3 percent drop in its own intelligence metric for Opus 4.6 and 4.7 and reverted on April 20. Twenty-five words is a tight enough budget that the model is being told to skip the reasoning it would otherwise narrate, and it seems that narration was doing work.
What the outside benchmark could and could not see
Margin Lab runs Claude Code every day against a curated, contamination-resistant subset of SWE-Bench-Pro, 50 instances per run, using the stock CLI rather than a custom scaffold so that the results reflect what a user would get. It models pass rate as a Bernoulli variable and flags statistically significant drops over daily, weekly and monthly windows. That design is exactly right for this failure mode, because it evaluates the product as shipped, including the system prompt, the caching behaviour and the reasoning default.
It also has limits the postmortem makes clear. At 50 instances a day the 95 percent interval on a single day's pass rate is wide, so a 3 percent effect from the verbosity prompt would be invisible day to day and only emerge over weeks. And the tracker is now collecting a fresh baseline for Opus 5.0 with degradation detection paused, which is the honest thing to do after a model change but leaves a gap in exactly the period you would most want coverage.
Why product-layer changes need an eval gate
We do not read this as carelessness. Each of the three changes had a reasonable motive, latency, cost and readability. The lesson we take is that a coding agent's behaviour is a function of the whole stack around the model, and only the weights were being treated as something that needed to pass an evaluation before release. A system prompt edit is a change to the policy. A caching rule that decides what context the model sees is a change to the policy. A default effort setting is a change to the policy. They should all clear the same bar.
The postmortem's own detection story supports this. The caching bug was masked by multiple internal experiments and surfaced through the /feedback command and a model-driven code review. A benchmark that ran the product end to end after every release would have caught the March 4 effort change within a day and the April 16 prompt change within a week. Margin Lab's setup is a proof that such a benchmark can be run from the outside for the cost of 50 SWE-Bench-Pro instances a day.
What we would do at a smaller lab
If you ship an agent product layer, the cheapest version of the fix is a fixed task suite that runs against the released binary, with the exact system prompt and defaults users get, on every release tag. Store pass rates with confidence intervals. Block releases that fall below the previous tag's lower bound. This is what we do for model checkpoints and it is strange that we do not do it for the code around them.
The second thing we would do is log the effective configuration with every session, the effort level, prompt version and caching mode, so that a user's complaint can be joined against what actually ran. The 6,852-session analysis in issue #42796 was heroic and it still pinned the blame on the wrong change, because the reporter could only see the surface of the product. Anthropic reset usage limits for all subscribers on April 23, which addresses the money. The trust question is whether the next change to the product layer gets an eval before it reaches users.
Sources
From the foundation