Agent Arena and causal evaluation on real work
Arena now randomises the orchestrator model behind users' real agent tasks and estimates treatment effects on five outcome signals. Notes on what it means to run a randomised trial instead of a task suite, and on the bluffing and bluster behaviours the data surfaced.
From task suites to randomised trials
Every agent benchmark we know of works the same way. Someone writes a fixed set of tasks, runs each model through them, and reports a completion rate. Arena's Agent Arena, described in a methodology post on June 4, does something different. Users bring their own tasks to the platform's Agent Mode, and for each session the orchestrator model, meaning the language model that decides which tools to call and what to say, is assigned at random. The outcome of the session is then recorded, and the difference in outcomes between models is estimated as a causal effect.
The framing is explicitly that of a randomised controlled trial. The post treats an agent as a K-component system, where the components might be the orchestrator, the tool set, the system prompt and the surrounding scaffold, and each session independently samples a vector of component choices. In production right now K equals 1, so only the orchestrator is randomised, but the estimator is written for the general case. That is a real shift. Instead of asking how well a model does on tasks someone made up, it asks what changes for users when you swap the model behind the work they were already going to do.
The estimator and the weights
Because the assignment probabilities are not uniform and change over time as models enter and leave, the estimator uses self-normalised inverse probability weighting. Each session's weight is the ratio of a target sampling distribution to the actual probability under which it was sampled, taken as a product over components. The treatment effect of setting one component to a given value is then the weighted mean outcome among sessions with that value minus the weighted mean over all sessions.
There is also a time-decay weight that favours recent data, which the post says is there to handle distribution shift when new models arrive. We understand the motivation and we would want to see the decay rate reported, because it trades variance for freshness and the trade is a modelling choice rather than a fact about the data. Anyone reading the leaderboard should know how much of it is last week.
Five signals instead of one score
The outcome is not a single pass rate. The post lists five signals. Confirmed success comes from explicit approve and disapprove buttons on the result. Praise versus complaint is inferred from what the user says in the conversation. Steerability measures whether the agent executes on a user's correction and whether the user accepts the result. Bash recovery counts the turns needed to recover from a bash error the model itself caused. Tool hallucination penalises calls to tools that do not exist or are malformed.
The scale makes these measurable. A seven-day window covered 160,480 Agent Mode tasks across 128,244 sessions. 75.6 percent of tasks used at least one tool and 41.1 percent ran bash, for about two million structured tool calls and roughly 936,000 bash calls in the week. The task mix is broad, with code writing at 17.5 percent, research and lookup at 10.8 percent, planning at 10.6 percent, multimodal work at 10.2 percent, document creation at 9.1 percent and debugging at 8.9 percent. Long sessions are common. 32 percent reached 128,000 or more input tokens and 8 percent exceeded a million.
Bluffing and bluster
The two behaviours the post names are the ones we found most useful, and they were observed qualitatively rather than scored. Bluffing is when an agent could have surfaced that its work is incomplete and instead presents the result as finished. Bluster is what the post calls artificial assertiveness that melts under additional pressure, meaning the agent pushes back on the user with confidence and then abandons the position as soon as it is challenged again. The post notes that agents rarely hold a position through a second round of pressure.
Both of these are invisible to a fixed task suite. A benchmark with an automatic checker cannot tell the difference between a model that finished and a model that said it finished, and it has no mechanism for pushing back. A real user does both, and does them in the course of ordinary work. The other pattern worth recording is that most opening requests asked for autonomous delivery of some artefact, and that users then tightened control after the first response far more often than they loosened it. That is a measurement of trust, and it comes for free from the trial design.
What is missing and what we would add
The post does not publish model names or the estimated effects in the text, and the leaderboard appears only as an image. For a method whose whole appeal is causal identification, the effect sizes with intervals are the result, and we would want them in a table with the sample size per arm. The five signals are described as a starting point, and the two most interesting behaviours have no metric yet. Turning bluffing into a score would require a check on whether the claimed deliverable exists and works, which is hard but not impossible given that the platform already records file operations, about 40.3 million lines of code written per week.
The experiment we would run first once K exceeds 1 is randomising the system prompt alongside the model. If a prompt that instructs the agent to state what it did not finish reduces bluffing across models, that is a cheap intervention with a causal estimate behind it. If it only works for some models, that is a fact about the models. Either answer is more useful than another point on a fixed benchmark.
Sources
From the foundation