2,778 researchers, 50 percent by 2047: reading the AI Impacts survey
The largest survey of published AI researchers to date moved the median forecast for high-level machine intelligence 13 years earlier in a single cycle. Reading notes on what the numbers say, how much the wording moves them, and what a survey of practitioners can and cannot tell us.
Who was asked
Katja Grace and colleagues at AI Impacts posted the results of their 2023 survey on January 5. They contacted 20,066 researchers who had published at top AI venues, reached 18,459 working email addresses, and got 2,778 responses, a 15 percent response rate, with 95 percent of those completing the survey. That is the largest sample of this kind anyone has collected, and it is the third cycle, so there are comparable numbers from 2022 to set it against.
The design carries over from earlier rounds. Respondents get one of several framings for the big questions, and the framings are randomised, so the survey can measure how much the wording moves the answer. That turns out to be the most useful feature of the whole exercise.
What moved
The headline is the forecast for high-level machine intelligence, defined as unaided machines accomplishing every task better and more cheaply than human workers. The aggregate median for a 50 percent chance is 2047. In the 2022 survey it was 2060. Thirteen years vanished from the median in one year, and that year contained the release of ChatGPT and GPT-4.
The companion question, full automation of labour, asks when for any occupation machines could be built to do the work better and more cheaply. The median for that moved from 2164 to 2116, 48 years earlier. Respondents were also asked about 32 specific tasks that had been in the 2022 survey, and for most of them the 2023 forecast came in sooner. Whatever one thinks of the absolute numbers, the direction is consistent across the whole instrument.
The framing gap
Here is the part we keep coming back to. The two big questions describe almost the same event. Every task done better and cheaper by machines, versus every occupation automatable. Yet the medians are 69 years apart. Respondents who got the occupation framing consistently gave much longer timelines than those who got the task framing. The survey authors flag this openly as a framing effect and do not try to reconcile it.
The extinction questions show the same sensitivity. Three versions were asked. On the general question about advanced AI leading to outcomes as bad as human extinction, the median was 5 percent and the mean 16.2 percent. On a version specifically about loss of control, the median was 10 percent. On a version with a hundred-year window, the median was back to 5 percent. Depending on which version you count, between 38 and 51 percent of respondents put the probability at 10 percent or higher.
So the honest summary of the extinction result is a range, and the range is wide. It is true that a large fraction of published AI researchers assign non-trivial odds to catastrophe. It is also true that a reasonable change in wording moves that fraction by 13 points. Anyone quoting a single figure from this survey is choosing a framing, whether they say so or not.
What practitioners know and do not
The authors are careful about the limits, and we want to restate them because the numbers will circulate without the caveats. Publishing at NeurIPS or ICML makes someone an expert in building systems. It does not make them an expert in forecasting, and there is no evidence in the survey that respondents with better calibration on near-term questions gave different long-term answers. The 13-year shift is the field updating on one year of visible progress, which is a fact about the field's mood as much as about the future.
What the survey does measure well is the distribution of belief inside the community that builds these systems. That is worth knowing on its own terms. If 68.3 percent think good outcomes are more likely than bad, and 48 percent of that optimistic group still give at least 5 percent to extinction-level outcomes, then the people doing the work hold a bimodal view of their own project, and that shapes what they choose to build and what they will accept in the way of oversight.
A worked example of how to use the result. Suppose a policy body wants to know whether researchers think a given capability is ten years away. The survey gives a median with a framing-dependent spread. The right reading is to take the spread as the answer. A capability whose median moves by decades under rewording is one the field has no settled view on, and a policy that assumes a settled view is built on sand.
What we would change next time
The next cycle should include calibration questions with resolvable short-term answers, so that responses on 2047 can be weighted by how the same person did on 2025. It should also ask the same respondent both the task and occupation framings and report the within-person gap, which would tell us whether the 69 years is a difference between people or a difference in how one person reads two sentences.
The thing we would not change is the randomised framing itself. It is the reason we know the numbers are soft, and a survey that reported one clean median would be less informative, whatever the median was.
Sources
From the foundation