Two reports, two weeks apart

Two pieces of evidence about AI adoption landed this month and they appear to point in opposite directions. The first is a report from MIT's NANDA initiative, The GenAI Divide: State of AI in Business 2025, which Fortune summarised under the figure that 95 percent of enterprise generative AI pilots are failing to produce measurable P&L impact. The second is a working paper from Erik Brynjolfsson, Bharat Chandar and Ruyu Chen at the Stanford Digital Economy Lab, which uses ADP payroll records covering millions of workers and finds that employment for 22 to 25 year olds in the occupations most exposed to AI has fallen while employment for older workers in the same occupations has held or grown.

If almost nothing works, why is anyone's headcount moving. If headcount is moving, why do pilots fail. We have spent the past week with both documents and we think the contradiction is mostly an artefact of what each one measures. The MIT report counts projects at the firm level. The Stanford paper counts people at the occupation level. Adoption that substitutes for a discrete task shows up in the second and barely registers in the first.

What the MIT report actually counted

We should be clear about what we have read. The Fortune summary reports that the study, led by Aditya Challapally, rests on 150 interviews with leaders, a survey of 350 employees, and an analysis of 300 public AI deployments. Its headline is that only about 5 percent of pilots reach rapid revenue acceleration. The detail underneath the headline is more useful than the headline. Purchased tools from vendors succeeded about two thirds of the time. Internal builds succeeded about one third of the time. More than half of generative AI budgets went to sales and marketing tools, while the largest measured returns came from back office automation.

That pattern is a story about integration cost, not about model capability. A firm that buys a working tool for a narrow back office task gets a return. A firm that funds an internal build of a broad sales assistant mostly does not. The 95 percent figure aggregates both and reads as a verdict on the technology, when it is closer to a verdict on where budgets were pointed in 2024 and early 2025.

The workforce finding in the same report is the bridge to the Stanford paper. Fortune reports that firms were not conducting mass layoffs but were increasingly declining to backfill vacant roles, with the changes concentrated in work that had previously been outsourced. Not backfilling a role is invisible in a pilot success count. It is very visible in payroll data for the youngest cohort, because they are the people who would have filled it.

What the payroll data shows

The Stanford paper is descriptive by design. The authors say plainly that they are documenting patterns rather than estimating causal effects, and the six facts are organised around that caution. The first fact is the one the title refers to. Employment for workers aged 22 to 25 in the most AI-exposed occupations, with software developers and customer service representatives as the running examples, has declined relative to its late 2022 level, while experienced workers in the same occupations and workers of all ages in low exposure occupations such as nursing aides have grown.

The second fact separates the young cohort's stagnation from the overall market. Total employment in the ADP sample kept growing, but for 22 to 25 year olds in the most exposed quintiles employment fell by about 6 percent from late 2022, against a 6 to 9 percent rise for older workers over the same period. The fourth fact adds firm-by-time fixed effects to absorb shocks that hit a whole firm regardless of AI exposure, such as interest rates, and the relative decline for the young exposed group survives at roughly 15 log points. The fifth fact is that the adjustment shows up in employment counts rather than in pay, which the authors read as either wage stickiness or offsetting effects.

The third fact is the one that ties the two reports together. The authors split occupations by whether observed Claude usage in that occupation's tasks is mostly automative or mostly augmentative, using the Anthropic Economic Index classification. Entry level employment fell in the occupations where AI use is automative. In the occupations where AI use is most augmentative, young worker employment was flat or grew. Substitution shows up as a missing hire. Augmentation does not.

The same curve from both ends

Put the two findings side by side and the shape is consistent. Where a task is cleanly substitutable, the adoption already happened and it happened quietly. A support queue that a model handles, a first draft that a junior would have written, a data pull that used to be a task for a new analyst. None of these required a pilot to succeed in the sense the MIT report measures, because the firm did not need to integrate anything into a workflow. It simply stopped needing the person at the bottom of the workflow.

Where the value depends on integration, adoption has stalled, and that is what the pilot failure rate is picking up. The broad internal builds that absorb most of the budget are attempts to insert a model into a process that was designed around people passing context to each other. Those projects fail for the usual reasons enterprise software projects fail, and the model is the least of it.

What we would want measured next

The obvious gap is that nobody has yet joined the two datasets. The MIT report has firm-level project outcomes and the Stanford paper has occupation-level hiring, and the question that matters is whether the firms not backfilling junior roles are the same firms whose pilots are counted as failures. Our guess is that they largely are, and that a firm can simultaneously report no P&L impact from its generative AI programme and stop hiring the people whose tasks the models absorbed, because the accounting for the second effect sits in a different budget line from the first.

The other thing we would want is a repeat of the Stanford exercise in a year. The authors are careful to say that some of the trends predate generative AI adoption and that controlling for education attenuates the results. If the gap between young and older workers in exposed occupations keeps widening while pilot success rates stay flat, the substitution story gets stronger. If pilot success rates rise as integration tooling matures and the hiring gap closes, the story was about a transition rather than a replacement. Either would be more informative than the current pair of headlines.

Sources

  1. Fortune, MIT report: 95% of generative AI pilots at companies are failing (Sheryl Estrada, August 18, 2025)
  2. Brynjolfsson, Chandar and Chen, Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence (Stanford Digital Economy Lab)