The framing everyone quoted

Karpathy's talk on Tuesday has three numbered eras. Software 1.0 is code a person writes. Software 2.0 is neural network weights. Software 3.0 is the prompt, in English, that programs a language model. His claim is that 3.0 is eating 1.0 and 2.0 and that a large amount of existing software will be rewritten. He also runs the analogies that have been repeated all week, LLMs as utilities like the electricity grid, as fabs with enormous capital costs, as operating systems, and as timeshared mainframes in a 1960s moment before the personal computer.

We want to spend our words on the parts that got less airtime, because they are the ones with design consequences. The era framing tells you where we are. The rest of the talk tells you what to build.

Jagged intelligence

The first of Karpathy's psychological observations is that model competence is uneven in a way human competence is not. The example he uses is the model that solves a hard problem and then says 9.11 is larger than 9.9. His phrasing is that some things work extremely well while some fail catastrophically. In people, skills develop in a correlated way, so a person who can do the hard thing can usually do the easy one. Models offer no such guarantee.

The design consequence is that you cannot infer reliability on one task from success on a related one. Every task needs its own check. That sounds obvious written down, and it is routinely violated, because the human instinct is to extend trust from a demonstrated success to the neighbourhood around it.

Anterograde amnesia

The second observation is that a model is like a coworker with anterograde amnesia. It has the context window and nothing else. Whatever it worked out with you yesterday is gone today unless you put it back in the prompt. Karpathy's suggested direction is what he calls system prompt learning, the model accumulating durable knowledge in its instructions rather than only in its weights.

This is the observation we think the industry has under-weighted. Most of the friction people report with assistants is a memory problem dressed up as a capability problem. The model can do the task. It cannot remember that you told it how you want the task done. Every tool that stores project notes, coding conventions, or a file describing the repository is an attempt to work around the same deficit, and the reason they all look alike is that they are all patching the same hole.

The autonomy slider

The design principle that ties the talk together is partial autonomy with a slider. Karpathy's examples are concrete. In Cursor the slider runs from Tab completion to Cmd K edits to Cmd L chat to Cmd we agent mode. In Perplexity it runs from search to research to deep research. In Tesla Autopilot it runs across levels. The user picks how much to hand over for this task, and the product has to work at every setting.

He pairs the slider with what he calls the generation and verification loop, and the point is speed. The model generates, the person verifies, and the productivity of the whole system is set by how fast the verification step can go. That argues for interfaces that make checking easy and for keeping the model on a tight leash, his phrase, rather than letting it run for an hour and dumping a result. His own report on vibe coding fits this. It went well while the code was local and fell apart at deployment, where the web stack is fragmented across services designed for human experts, and a person had to click through all of it.

Our favourite line was the distinction between a demo and a product. A demo is works.any(). A product is works.all(). The gap between those two is where most of the engineering lives, and it is invisible in a launch video.

Build for agents

The last section argues that there is now a third consumer of digital information alongside humans with GUIs and computers with APIs, and that this consumer, the agent, is like a human in what it wants but a computer in how it reads. HTML built for people is a poor fit. Karpathy points to llms.txt as a way of telling agents what a site is, and to context builders like Gitingest that flatten a repository into something a model can read in one pass.

What we take from the talk as a whole is that the interesting work in the next few years is at the seams. Memory across sessions, verification interfaces fast enough to keep a human in the loop, and documentation written for a reader that cannot click. None of that requires a bigger model. All of it requires deciding, task by task, where the slider should sit, and we would like to see teams publish that decision alongside their accuracy numbers.

Sources

  1. Latent Space, Software 3.0 (transcript and notes on Karpathy's talk at YC AI Startup School, June 17, 2025)