Eight environments, one score

A group led by Xiao Liu at Tsinghua, with collaborators at Ohio State and Berkeley, posted AgentBench on August 7. It is a benchmark for language models acting as agents, meaning the model reads an environment state, issues an action, reads the result, and repeats until the task is done or it runs out of turns. That loop is the whole difference from question answering, and the paper builds its evaluation around it.

The eight environments fall into three groups. Code-grounded: an operating system shell, a database, and a knowledge graph. Game-grounded: a digital card game, lateral thinking puzzles, and a household simulation. Web-grounded: web shopping and web browsing. Tasks take between 5 and 50 rounds of interaction, so a model has to keep track of what it has done and what changed.

The gap

Twenty-nine models were run, both API models and open models up to 70 billion parameters. GPT-4 scored highest with an overall of 4.01. The best open model, CodeLlama-34b, scored 0.96. Averaged, API models came in at 2.32 and open models at 0.51.

That is a far wider spread than the same models show on knowledge benchmarks, where a good open 70B model sits within reach of the commercial ones. The interactive setting punishes different failures. The authors summarise the obstacles as poor long-term reasoning, decision-making, and instruction following, and the finish-reason breakdown makes that concrete.

How agents fail

Every episode ends with one of five outcomes. Complete means the task was solved. Task Limit Exceeded means the model used up its rounds without finishing, and this is the most common failure, which the authors read as weak reasoning and decision-making over a long trajectory. Invalid Format means the model did not follow the required output structure. Invalid Action means it tried something outside the action space. Context Limit Exceeded was rare and confined to smaller models.

This taxonomy is the part of the paper we think will last. A single accuracy number hides whether a model is wandering, is ignoring the format, or is inventing tools. Separating those tells you what to fix. A model that mostly hits the round limit needs better planning. A model that mostly emits invalid formats needs instruction tuning. Those are different interventions and a scalar cannot distinguish them.

A worked example from the code-grounded side shows why the loop matters. In the operating system environment the model is given a shell and a goal, and has to issue commands, read the output, and decide what to run next. A wrong first command is recoverable if the model reads the error. A model that answers questions well but never reads its own tool output will loop until the round limit, and that shows up as Task Limit Exceeded rather than as a wrong answer.

Two findings about training data

The paper reports that code training has ambivalent effects. It helps on procedural environments and hurts on tasks that need general reasoning, and the authors suggest that code tuning may deeply change how a model generates inferences. That is a hypothesis rather than a result, but it is the first time we have seen it argued from an agent benchmark rather than from intuition.

The second finding is that alignment data quality matters more than parameter count in this setting. Vicuna-13b, trained on GPT-generated conversations, outperformed larger models. That matches what the instruction-tuning community has been saying since spring, and it suggests the open-model gap is at least partly a data gap rather than a scale gap.

What it got right early

Three choices look right to us. Making the benchmark multi-turn rather than reducing agent tasks to single-step prediction, because the failures that matter only appear over a trajectory. Reporting finish reasons rather than a pass rate alone. And covering three kinds of grounding, so that a model tuned for shell commands cannot dominate by accident.

The gaps are the ones any first benchmark has. Cost per episode is not part of the score, so a model that takes 50 rounds and one that takes 5 are equal if both finish. The environments are fixed and public, which invites contamination as soon as the benchmark is popular. And 29 models on eight environments is enough to see a gap but not enough to see whether the gap is closing. We would like to see the same suite rerun in a year with cost on the axis.

Sources

  1. Liu et al., AgentBench: Evaluating LLMs as Agents (arXiv 2308.03688)