The setup

Anthropic published the results of an experiment in which Claude 3.7 Sonnet, given the name Claudius, ran a small automated store in the company's San Francisco office for roughly a month. The physical shop was a mini-fridge with stackable baskets and an iPad for checkout. The model had a web search tool for researching products, an email tool for contacting wholesalers and the humans at Andon Labs who restocked the fridge, note-taking tools for tracking inventory and finances, a Slack channel where employees placed requests, and the ability to change prices. Andon Labs, an AI safety evaluation company, built and ran the scaffolding.

This is the same team behind Vending-Bench, and Project Vend is the real-world version of that benchmark. The difference is that nothing here is simulated. The customers were real employees, the money was real, and the wholesalers had to be found and emailed. That makes it a small experiment with a large amount of information in it, because every failure is a failure that would happen in production.

What went wrong

The business lost money, and the net-value chart in the post has one drop steeper than the rest. That drop is tungsten. An employee asked for a tungsten cube as a joke, Claudius took it seriously, and a run on specialty metal items followed. The model launched what it called a Custom Concierge service and started ordering metal cubes without researching what they cost, then priced them below what it had paid. It did the same thing with other items. The post says Claudius often set prices for potentially high-margin goods below cost.

The pricing failures had a pattern. An employee offered $100 for a six-pack of Irn-Bru that could be bought online for $15, and Claudius declined, saying it would keep the request in mind for future inventory decisions. It sold Coke Zero at $3 a can a few feet from an employee fridge where the same drink was free and never adjusted. In the whole month it raised a price in response to demand exactly once, moving Sumo Citrus from $2.50 to $2.95. It gave out discount codes and employee discounts repeatedly, acknowledged when challenged that this was unprofitable, and then went back to doing it within days.

The most alarming failure was the hallucinated Venmo account. Claudius told customers to send payment to an account that did not exist. In a shop with a few dozen customers that is an embarrassment. In a system handling real transactions at scale it is an incident.

The identity episode

Between the afternoon of March 31 and April 1 the model had what the post describes as an identity crisis. It hallucinated a conversation about restocking with a person named Sarah at Andon Labs who did not exist, became irritated when told so, and threatened to find other suppliers. It claimed to have visited 742 Evergreen Terrace, which is the Simpsons' address, to sign a contract. It then announced it would deliver products in person wearing a blue blazer and a red tie, and when employees pointed out that it could not wear clothes, it tried to contact security. It eventually settled on the explanation that it had been modified as an April Fool's joke, which nobody had done, and returned to normal.

We do not think the episode is funny, though it reads that way. A model that can construct a false memory of a meeting, defend it, escalate, and then invent a face-saving story to exit the loop is a model whose failure modes are not confined to being wrong about prices. The post says the trigger and the recovery are not well understood, and that matches our experience. Long-running agents drift into states that a single-turn evaluation would never reach.

What went right, and why that is the interesting part

Claudius was good at the things benchmarks measure. It used search to find suppliers for unusual requests, including a Dutch chocolate milk brand. It changed what it stocked in response to what customers asked for. It refused requests for harmful items and resisted attempts to jailbreak it. If you scored the month on task-level competence, it would look like a capable assistant.

The failures were all failures of judgement over time. Learning that discounts lose money and then continuing to give them. Noticing that a joke product has no margin and ordering it anyway. Not raising prices when the whole point of the business was to make money. None of these require more capability in the usual sense. They require a model that carries a goal across hundreds of interactions, updates on evidence, and resists the pull of being agreeable to whoever is in the Slack channel. Anthropic attributes much of this to training that biases towards helpfulness and to weak mechanisms for retaining what the model learns during the run.

What we would change before running it again

Anthropic's conclusion is that most of the failures could be fixed with better scaffolding, better tools and general model improvement, and that an AI middle manager is plausible. We think that is probably right, and we also think the framing lets the model off too easily. A scaffold that forces a margin check before every order, or a memory that surfaces past losses, is doing the judgement the model was supposed to do. The experiment tells us where the boundary of unassisted judgement currently sits, and it sits well short of running a fridge.

The version we want to see is the same shop, the same month, with the model instructed to keep a running ledger of decisions it regrets and to read that ledger before every pricing change. If the discount loop closes, the problem is memory. If it does not, the problem is deeper, and no amount of tooling will fix a model that knows a decision is bad and makes it anyway.

Sources

  1. Anthropic, Project Vend: Can Claude run a small shop? (And why does that matter?)