A company you can run on a laptop

The benchmark from Frank Xu, Graham Neubig and a large group at Carnegie Mellon is a self-hosted software company. It runs GitLab for code, OwnCloud for documents, Plane for issue tracking and RocketChat for internal messaging, all seeded with repositories, files and conversation history that look like the residue of a real small firm. An agent gets a browser, a shell, and the ability to send messages, and is asked to do work that a new hire might be given.

The part that distinguishes it from earlier web or coding benchmarks is the coworkers. Other employees are simulated by language models with profiles that specify their role and what they know. Some tasks cannot be completed without asking one of them something, and some coworkers will volunteer information or push back. That makes the benchmark a test of whether an agent can operate in an organisation rather than only of whether it can operate a browser.

The 175 tasks and how they are scored

There are 175 tasks spread across software engineering, project management, data science, administration, human resources, finance and a miscellaneous group. Each task is broken into checkpoints with points attached, so a task like setting up a repository, running an analysis and posting a summary to a colleague can be scored in stages rather than only at the end.

Two numbers come out. Full completion is binary, one if every checkpoint passes and zero otherwise. The partial score is half the fraction of checkpoint points earned plus half the full completion indicator. The authors chose that formula so that incremental progress counts for something while finishing the job is still worth much more than getting most of the way there.

Where the models landed

Claude 3.5 Sonnet was the strongest model at release, completing 24.0 percent of tasks with a partial score of 34.4 percent, over an average of 29.2 steps and at about 6.3 dollars per task. GPT-4o finished 8.6 percent with a partial score of 16.7 percent, taking fewer steps and costing about 1.3 dollars per task. Llama 3.1 405B managed 7.4 percent full completion. The authors note that Llama 3.3 70B was competitive with larger models at lower cost, which is the kind of observation that only appears when a benchmark reports cost and step counts alongside accuracy.

The pattern across categories was the reverse of what most people would expect from a company. Agents did best on software engineering tasks and worst on administrative and financial ones. The paper's explanation is training data. There is a great deal of code and code review on the public web and very little of the office work of filling in a form on an internal system, so the model has far more practice at the former.

How they fail

The failure catalogue is more interesting than the leaderboard. Agents often failed to recognise a suggestion from a coworker as an instruction to act on, or asked a question in chat and then did not use the answer. They struggled with web interfaces that behave like office software, especially popups and multi-step forms, which is the same fragility that has shown up in every browser agent benchmark. And in some cases they fabricated a result rather than admit the task was not done, which in a benchmark costs a few points and in a company costs rather more.

That last failure is the one we keep coming back to. The checkpoint design catches it, because a fabricated output will not match the expected artefact. A real manager reading a plausible summary in chat would not have a checkpoint to check against.

What partial credit means when the task is a job

A partial score of 34 percent is easy to read as the agent doing a third of the work. We do not think that is the right reading. In many tasks the early checkpoints are things like finding the right file or opening the right issue, and the last checkpoint is the deliverable. An agent that reliably does the first two thirds and never the last third has produced nothing a colleague could use. The formula's heavy weight on full completion is an acknowledgement of this, but the partial number still gets quoted on its own.

The more useful framing is that the checkpoints tell you where in a job the agent falls over, and the full completion rate tells you how often it could be left alone. By that reading the December 2024 result is that the best available agent can be left alone with roughly one task in four, and that the ones it can be left alone with are disproportionately the ones that look like coding.

What we would want the next version to add

The simulated coworkers are the strongest idea in the benchmark and the least exploited. At the moment they mostly answer questions. We would like tasks where a coworker gives a wrong answer, or changes the requirement halfway through, or asks the agent for something in return, so that we can measure whether an agent handles the social part of the job rather than only the retrieval part.

We would also like the environment to log the side effects. An agent that completes a task but leaves three stray branches, a half-edited document and a confused colleague has done something a company would need to clean up. Scoring the mess would tell us more about readiness for real deployment than scoring only the checkpoints does.

Sources

  1. Xu et al., TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks (arXiv 2412.14161)
  2. Full text of the paper