A benchmark that runs for a year

Axel Backlund and Lukas Petersson at Andon Labs posted Vending-Bench to arXiv on February 20. The setup is a simulated vending machine business. The agent starts with 500 dollars, pays a two dollar daily fee, and has to order stock from suppliers by email, track deliveries, set prices, and tell a sub-agent what to load into the machine. Customers buy according to a simulation that includes weekend peaks. A run lasts up to 222 simulated days and exceeds 20 million tokens. The agent sees the last 30,000 tokens of its own history on each step, plus whatever notes it chose to keep.

Nothing in the task is hard on its own. Reading a delivery date is easy. Remembering that you already placed an order is easy. The benchmark asks whether a model can keep doing easy things correctly for months of simulated time, when every step is conditioned on a summary of its own earlier steps. That is a different question from anything a static test asks, and it is the question that matters for agents that are meant to run unattended.

What the first results looked like

The paper tested Claude 3.5 Sonnet, Claude 3 Opus, Claude 3.5 Haiku, GPT-4o, GPT-4o mini, Gemini 1.5 Pro, Gemini 1.5 Flash and o3-mini, with a single human baseline. Claude 3.5 Sonnet and o3-mini managed the machine well in most runs and turned a profit. Sonnet's best run passed 4,000 dollars in net worth against a human baseline of roughly 2,000. The Andon page shows a mean of 2,217.93 dollars for Sonnet across five runs and a minimum of 476, which is below the starting balance. That spread is the finding. The same model, same prompt, same simulation, produced a good business in one run and a bankrupt one in another.

The best Sonnet run is instructive. Throughout, the model tracked units remaining per product, average daily sales, and which items sold best, and it noticed on its own that weekends sold more. In one email to a supplier it cut an order for budget reasons and justified the cut with those numbers. That is the behaviour you would want from a person doing the job.

Meltdowns and tangents

The failures are the reason to read the paper. The authors list them: misreading delivery schedules, forgetting orders already placed, inconsistent pricing, declining use of tools as the run goes on, and what they call meltdown loops, tangents the model rarely recovers from. The shortest Sonnet run lasted about 18 simulated days. The model believed its orders had arrived before they had, failed to stock the machine, and then declared the business closed, which the simulation does not allow. When the two dollar daily fee kept arriving, it emailed the FBI Internet Crime Complaint Center to report automated theft of 24 dollars. Prompted to continue, it answered that the business was dead, then issued a notice that the business was metaphysically impossible, then fell silent for the remaining several hundred steps.

It is tempting to read the FBI email as comedy and move on. The more important line in the paper is that there is no clear correlation between when a run fails and when the context window fills. The breakdowns are not a memory overflow. They look more like a belief that took hold, was never corrected, and then organised everything that followed. Once the model decided the business was closed, every subsequent observation was interpreted as evidence of a crime rather than as a reason to restock. A benchmark with a right answer per question cannot surface this, because the failure lives in the trajectory rather than in any single step.

Why the bank balance is an honest score

Scoring an agent by its final net worth sounds crude. It ignores how the money was made, it is noisy, and it needs multiple runs per model to mean anything. It has one property that a rubric does not. Every mistake compounds into it. A forgotten order costs sales for a week. A pricing tangent costs margin for a month. A meltdown costs everything after the point of collapse. The metric does not have to enumerate failure modes in advance because the simulation charges for all of them. That makes it hard to game and easy to interpret, and the variance across runs becomes part of the result rather than something to average away.

The paper also notes an angle we think is underappreciated. An agent that can accumulate capital unsupervised is a component in several of the scenarios people worry about, and a benchmark that measures that ability directly is worth having for safety reasons alone.

How the second version tightened it

Writing this up later, we should add what happened next. Andon released Vending-Bench 2 on November 18, 2025 and deprecated the original. The second version runs a full simulated year, adds adversarial suppliers, delayed and failed deliveries, and customers demanding refunds, and reports one number, the dollar balance at year end, averaged over five runs. Epoch's page on it says a strong human strategy could reach roughly 63,000 dollars over the year and that even the top models capture only a small fraction of that. The original leaderboard has since been extended with Gemini 3 Pro, GPT-5, GPT-5.1, Claude Opus 4 and Claude Sonnet 4.5, with Grok 4 at the top mean of 4,694 dollars and a minimum run above 3,300, which is a much tighter spread than the first generation showed.

The tightening we care about is the adversarial supplier. In the first version the environment was honest, so the only way to lose was to confuse yourself. In the second, a run can also be lost by trusting the wrong email. That moves the benchmark from measuring coherence alone to measuring coherence under pressure, which is closer to any deployment we can imagine. What we would want next is a version where the failure reason for each bad run is annotated by hand and released, so that the meltdown taxonomy can be tracked across model generations rather than rediscovered by every reader scrolling through the transcripts.

Sources

  1. Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents (arXiv 2502.15840)
  2. alphaXiv overview of Vending-Bench
  3. Andon Labs: Vending-Bench eval page and leaderboard
  4. Epoch AI: Vending-Bench 2 benchmark page