What shipped

On October 22 Anthropic released an upgraded Claude 3.5 Sonnet and, with it, a public beta of something it calls computer use. The developer gives the model a screenshot, the model returns an action such as moving the cursor to a coordinate, clicking, or typing text, the developer executes it and sends back a new screenshot, and the loop repeats. Anthropic's description is that the model can look at a screen, move a cursor, click buttons and type, and the examples given are checking a spreadsheet, opening a browser, filling in forms.

This is different from every previous agent API in one specific way. It does not need the application to expose a tool. Function calling required someone to write a schema for each capability. A browser agent required a DOM. Computer use needs pixels, which means in principle it works on any software a person can see. That is why it is the first general-purpose one, and also why it is so slow.

The numbers, and the human range

The headline benchmark is OSWorld, a suite of real tasks on real operating systems. Claude 3.5 Sonnet scores 14.9 percent in the screenshot-only setting, against 7.8 percent for the next best system, and 22.0 percent when allowed more steps. Anthropic puts typical human performance at 70 to 75 percent. So the best available system fails roughly six tasks in seven, on a benchmark whose tasks a person completes without thinking.

The same model moved from 33.4 to 49.0 percent on SWE-bench Verified and from 62.6 to 69.2 on the retail split of TAU-bench, which are respectable jumps in domains where the model already had a tool interface. The contrast with OSWorld is the point. Give the model a shell and it is a useful junior engineer. Give it a screen and it is a beginner.

Why pixel control is slow and brittle

Anthropic's engineering post explains how the model was taught to click at all. It was trained to count pixels vertically and horizontally from screenshots to estimate where the cursor should go, and the training used only simple software such as a calculator and a text editor, with no internet access for safety reasons. The post says the model generalised from that quickly, and we believe it, but the method tells you the failure mode. A click is a regression to a coordinate on an image, and a few pixels of error on a small button is a miss.

The second problem is that the model sees the screen as a flipbook. Each screenshot is a still, so anything that happens between frames, a notification that appears and disappears, a loading spinner, a dropdown that closes on the next click, is invisible. Anthropic says outright that scrolling, dragging and zooming, which people do without effort, currently present challenges, and describes the whole capability as experimental, at times cumbersome and error-prone, and slow.

The failures in the company's own demos are the honest part of the release. In one, the model accidentally clicked the button that stopped a screen recording and lost the footage. In another, in the middle of a coding task, it opened a browser and started looking at photographs of Yellowstone National Park. Nobody at Anthropic asked it to. That is what a distracted agent looks like when the action space is the entire desktop, and it is a reminder that a wider action space is not free.

The gap between the demo and the deployment

The demos are compelling because a successful run looks like magic. A model reads a spreadsheet, opens a website, fills a form and submits it, and for thirty seconds the future has arrived. The benchmark says the run succeeds about one time in seven. Both facts are true, and the interesting work is in the difference.

For a demo you need one successful trajectory. For a deployment you need the failure cases to be cheap, and on a desktop they often are not. A form submitted with the wrong values, a file deleted, an email sent to the wrong person, are all one click away from any task, and the model cannot undo a click. Anthropic has built classifiers to detect when computer use is running and whether harm is occurring, and has added election-specific measures to steer the model away from posting on social media, registering domains, or interacting with government sites. The threat it flags most directly is prompt injection from the screen itself: text on a web page that the model reads as an instruction.

What we would want someone to try is to treat the screenshot loop as a fallback rather than the primary interface. Use tools where tools exist, and drop to pixels only for the software that has none. That hybrid would score better than either approach alone, and more to the point it would fail more gracefully. The interesting question for the next year is whether the pixel-only score climbs fast enough that nobody bothers, or whether 14.9 percent is what it looks like when a model learns to see a screen without being given a map of it.

Sources

  1. Anthropic, Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku (22 October 2024)
  2. Anthropic, Developing a computer use model