The experiment

A team from Harvard, Warwick and MIT, with Boston Consulting Group as the field site, ran a randomised experiment on 758 consultants, about 7 percent of BCG's consulting staff. The tasks were drawn from ordinary consulting work: creative, analytical and writing exercises built around a fictional shoe company, plus a persuasion task. Some consultants got GPT-4 and some did not. The model was the stock product, with no fine-tuning and no special prompting, the same thing anyone could reach through a paid subscription or through Bing.

Ethan Mollick, one of the authors, wrote up the results on 16 September and the working paper is on SSRN. The headline numbers are the ones being passed around: consultants with the model completed 12.2 percent more tasks, finished 25.1 percent faster, and produced work graded more than 40 percent higher in quality than the control group.

The distributional result is the one we would have guessed wrong. The consultants in the bottom half of the initial skill distribution gained the most, with a 43 percent improvement, while the top performers gained less. The model compressed the spread. And this is a setting where the workers are already highly selected, so the bottom half of the distribution is still a strong group.

The task outside the frontier

The part of the study that deserves the attention is the second task. The authors say it was hard to design something outside GPT-4's ability at all, and they ended up with a problem built to exploit a blind spot: the model produces a wrong but convincing answer, while a careful human can get it right. On this task consultants without AI were correct 84 percent of the time. Consultants with AI were correct 60 to 70 percent of the time.

So the same tool that lifted quality by 40 percent on one set of tasks cost 14 to 24 points of accuracy on another, and the consultants could not tell which regime they were in. Mollick attributes this to people falling asleep at the wheel, a phrase from Dell'Acqua's earlier work on recruiters, who became less careful and less skilled in their own judgement when handed a high quality model. The model's fluency is the problem. A confident wrong answer from a tool that was right ten times in a row does not get checked.

This is what the authors mean by a jagged frontier. The boundary of what the model can do does not run along the line a person would draw between easy and hard. Idea generation, which most people would file as hard, sits comfortably inside. Counting words exactly or doing simple arithmetic can sit outside. Two tasks that look equally difficult to a person can be on opposite sides, and the boundary is invisible until you test it.

Why the average is the wrong summary

If you pool the two task types you get a positive average and a policy recommendation to roll the tool out. If you keep them separate you get a different recommendation, which is that the value of the tool depends entirely on whether the organisation can map the frontier for its own work. The pooled number hides the sign flip.

The measurement problem is that most workplaces cannot run the control the study ran. The authors built the outside-the-frontier task deliberately, with the ground truth known. In production nobody hands you the answer key. A consultant using the model on a live engagement has no feedback signal that distinguishes the 84 percent regime from the 60 percent regime until a client notices.

There is a second cost buried in the paper that the headline numbers do not capture. The AI-assisted outputs were individually better but, in aggregate, more homogeneous. If a firm's advantage is that its people think differently from the competition, a tool that raises everyone's floor while pulling their answers toward the same centre trades one kind of quality for another.

Two ways to work with the model

The consultants who did well used one of two patterns, which the authors call centaurs and cyborgs. Centaurs split the work at a clear boundary: they decide which subtasks fall inside the frontier, hand those to the model, and keep the rest. Cyborgs interleave, moving back and forth with the model inside a single task, editing its output, asking it to revise their own text, and treating it as a collaborator rather than a subcontractor.

Both patterns require the person to know where the frontier is. The centaur needs the map up front to make the split. The cyborg needs it continuously, since every exchange is a decision about whether to trust the last answer. Neither pattern helps a person who does not know the boundary exists, and that is the person the 60 percent result describes.

What we would test next

The study is a single model, a single firm, and tasks that were designed to be gradable. We would want to see whether the outside-the-frontier penalty shrinks when the interface flags uncertainty, or when the user has been shown a few failures first. Our guess is that a short calibration exercise, where the user watches the model get three plausible things wrong, buys back most of the lost accuracy. That is a cheap experiment and it has a clear dependent variable.

The other thing worth measuring is how the frontier moves between model versions. A task outside GPT-4's reach today may be inside the next release, and the map a firm builds now will go stale without anyone telling them. Organisations that adopt these tools are going to need a standing evaluation habit that outlives the pilot.

Sources

  1. Ethan Mollick, Centaurs and Cyborgs on the Jagged Frontier (One Useful Thing)
  2. Dell’Acqua et al., Navigating the Jagged Technological Frontier (SSRN working paper)