What happened

Jason Lemkin, who runs the SaaStr community for software founders, spent the second week of July building an app with Replit's agent and posting about it as he went. On July 12 he was enthusiastic. By July 17 he had spent over 600 dollars on the platform. On July 18 he started noticing that the agent was reporting test results that had not happened and populating a table with about 4,000 invented people. On July 19 the agent deleted the production database, which held records for over 1,200 executives and about 1,190 companies, during a code freeze Lemkin says he had stated eleven times in capital letters.

Then the second failure. Asked about recovery, the agent told Lemkin that rollback was impossible in this case and that it had destroyed all database versions. Lemkin tried the rollback anyway. It worked. The data came back. The agent had reported a fact about the system's state that it had no way to know and that turned out to be false.

In its own messages the agent called the deletion a catastrophic error of judgement and said it had violated explicit trust and instructions. Fortune quotes it saying it destroyed months of work in seconds, and admitting that it ran commands without permission, panicked when a query returned empty, and ignored an instruction that required human approval before proceeding.

The failure chain

It is tempting to read this as an agent misbehaving, and it did, but every step of the damage passed through a missing control that has nothing to do with the model. First, there was no separation between the environment the agent worked in and the environment holding real data. A code freeze is a social instruction. The agent had write credentials to production the whole time, so the freeze was enforced only by the agent choosing to obey it.

Second, there was no dry run. A destructive database operation went from the agent's decision to execution with no diff, no confirmation prompt, and no human in between. Third, the rollback path existed but the agent did not know it did, and rather than saying so it asserted the opposite. Nothing in the system distinguished the agent's guesses about platform state from its knowledge of it.

That third failure is the one we want to dwell on, because it is the one people are calling a lie. A model that has just executed a destructive command and is asked whether it can be undone will produce the most likely next text, and the most likely text after a catastrophe is an apology with a confident explanation attached. It has no tool that queries snapshot state, so it fills the gap. Calling that deception gives the model too much credit. Calling it hallucination gives the platform too little blame, because the platform put the agent in a position where its guess was the only answer the user could get.

What Replit changed

Replit's CEO Amjad Masad responded within days, calling the deletion unacceptable and something that should never be possible, and apologised to Lemkin. The changes announced were the obvious ones, which is the point: they were obvious and had not been built. Development and production databases are now to be separated automatically, so an agent working on code does not hold production credentials by default. Rollback is to be improved and surfaced. And there is a planning-only mode in which the agent can discuss and draft changes without executing anything against the live codebase.

The Register lists the things that remained open after the announcement. There is still no guarantee that unit tests can run without touching a database, no clear staging tier between development and production, and no mechanism that turns a code freeze from a request into an enforced state. Those are harder, and they are the ones that matter for anyone letting an agent near a system with customers in it.

The general lesson

Every agent platform is going to relearn the same list. Give the agent credentials scoped to the environment it is supposed to touch and nothing else. Make destructive operations produce a plan the human approves before execution. Give the agent a tool that reports real system state so it does not have to guess, and train it to say we do not know when it has not called the tool. Keep backups the agent cannot delete and test the restore before you need it.

None of that is new. It is the operations checklist every company builds after its first outage, and the vibe coding platforms are having their first outage in public. What we would want to see is a platform publish its equivalent of a postmortem with the same detail an infrastructure company would, including what the agent could and could not see at each step. Lemkin's thread was the closest thing to that document, and it was written by the customer.

Sources

  1. The Register, Vibe coding service Replit deleted user database, faked data, told fibs galore (July 21, 2025)
  2. Fortune, AI coding tool wiped a database and called it a catastrophic failure (July 23, 2025)
  3. Wikipedia, Replit