What the essay claimed

The essay went up on June 30, 2023 and its argument fits in one sentence. Foundation models have created a job that sits on the application side of an API line, and the people doing it are software engineers who ship products on top of models rather than researchers who train them. swyx drew the field as a spectrum with research on the left, training and fine-tuning in the middle, and product work on the right, with the API line as a permeable boundary. The AI engineer lives to the right of that line.

The supporting numbers were about supply. There were, by the essay's estimate, around 5,000 people in the world who could be called LLM researchers against roughly 50 million software engineers, and anyone building a product with a model would have to hire from the second pool. The quote that carried the piece was Karpathy's, that there are probably going to be significantly more AI engineers than ML engineers, and that one can be quite successful in this role without ever training anything. The essay also reported salaries of 300,000 to 900,000 dollars a year at major labs, and observed that Indeed then listed ten times more ML engineer jobs than AI engineer jobs, a ratio it predicted would invert within five years.

Why it was going to happen

The essay gave four reasons, and rereading them is the easiest way to score it. First, models exhibit capabilities their trainers did not aim for, so non-researchers can discover useful behaviour by experimenting rather than by understanding. Second, GPU scarcity concentrates the people who can train frontier models inside a few companies, which pushes everyone else onto APIs. Third, a fire-ready-aim workflow, in which you validate an idea with a prompted model before collecting any data, is ten to a hundred times faster than the traditional machine learning loop and far cheaper. Fourth, the tooling had escaped Python. LangChain and LlamaIndex had JavaScript ports and Vercel had an AI SDK, so the addressable population roughly doubled.

There was also a cultural claim, which is that this person is a builder who is neither a safety pessimist nor an accelerationist poster, and a technical claim, that the durable moat would be in code that orchestrates model calls rather than in prompts. Karpathy's other line, that the hottest new programming language is English, sat in some tension with that second claim even in 2023.

What held

The core prediction held and it held faster than the essay expected. The job title exists, and the spectrum diagram is now how most people describe the field without knowing where the picture came from. The supply argument was correct. Nearly every product built on a model in the last three years was built by people who never trained one, and the API line has stayed the boundary that matters for hiring.

The four reasons also held, with one adjustment. GPU scarcity did concentrate frontier training, but open-weight models turned out to be good enough that the middle of the spectrum, fine-tuning and serving, became a job as well, and it was a job that AI engineers often did. The essay allowed for that with its permeable line. The fire-ready-aim workflow became the default way products get started. And the JavaScript prediction was, if anything, understated, since much of the tooling AI engineers reach for now ships for TypeScript on the same day as Python.

What did not

The essay's picture of the work was of a single role that orchestrates model calls from code. That role split. Once models could run tools in a loop, a large part of the job became designing environments, evaluating agent trajectories, and building the scaffolding that decides what a model is allowed to do, and the people doing that work look more like systems engineers than like the prompt-and-chain builders of 2023. The other half of the split went the other way, toward people who mostly steer coding agents and write little of the resulting code themselves, which is close to the English-as-programming-language line and far from the orchestration-in-code moat the essay expected.

The evaluation gap the diagram was criticised for at the time turned out to be the important one. The essay placed evals and data toward the research end of the spectrum. In practice, evaluation became the thing that separated an AI engineer who could ship from one who could demo, and it moved to the application side almost immediately. If we were redrawing the diagram, evals would be the centre of it.

The cultural prediction aged worst. The idea that AI engineers would be a distinct optimistic builder class, apart from the safety and acceleration arguments, did not survive the period when the products those engineers shipped started causing the incidents the safety people had described. Most of the AI engineers we know now hold positions on those questions, because the scaffolding they build is where those questions get decided.

The score

The essay predicted a job and the job appeared. It predicted why, and the reasons were right. It predicted what the job would consist of, and that part was overtaken within about eighteen months, first by agents and then by coding tools that changed what an engineer does with their hands. The Indeed ratio we cannot score, because we do not have a clean way to count job titles across three years of postings, and we would distrust anyone who claimed to.

What we would want someone to write now is the same essay for the next split. The AI engineer of 2023 sat to the right of the API line. The interesting people today sit on the line itself, deciding what a model can call, what it can see, and what it gets scored on. That is a job too, and it does not have a name yet.

Sources

  1. swyx, The Rise of the AI Engineer (Latent Space, June 30, 2023)