What happened

Around April 14, Cursor users who switched from a desktop to a laptop started getting logged out of the other machine. No notice, no changelog entry. Some of them emailed support to ask whether this was deliberate, and the reply said it was expected behaviour under a new login policy. That reply came from an AI agent answering the support inbox, and the policy it described did not exist.

The story travelled fast because it came from the company's own support channel. The user who wrote up the incident on Reddit, later reposted to Hacker News, said dozens of people publicly cancelled their subscriptions within hours, including the author. The original Reddit thread was locked and then removed, which read to onlookers as silence rather than a fix. By the time the actual cause was known, the cancellations had already happened.

The bug underneath

Two days later, Cursor cofounder Michael Truell replied on Hacker News. The logouts were a race condition that showed up on very slow internet connections, where a burst of unneeded sessions crowded out the real one. The fix had been rolled out. He confirmed that AI-assisted responses were used as the first filter for email support, said those responses would now be clearly labelled as such, and said the user who wrote the post had been fully refunded. Wikipedia's account of the incident, citing Ars Technica's April 21 report, describes the company apologising and refunding affected users.

So the technical fault was small and ordinary. A session bug that affects users on slow connections is the kind of thing that gets a hotfix and a one-line release note. What turned it into a story was the confident, wrong explanation, and that explanation was produced by a system whose job was to answer questions it had no grounded way to answer.

Why a good model was not enough

The bot did what a language model does when asked to explain something it has no record of. It produced the most plausible account. A sudden logout behaviour, a user asking whether it is intended, and the training distribution is full of support emails that say yes, this is expected behaviour under our updated policy. Nothing in the model's inputs told it that Cursor had shipped no such policy, because nothing in its inputs was Cursor's actual policy or Cursor's actual incident log.

This is the part we want to press on, because the fix is often described as a better model, and it is not a model quality problem. A stronger model would still have had no fact to check against. Front-line support answers about product behaviour need to be grounded in a source the company controls, a policy document, a status page, a known-issues list, and the agent needs to be able to say that it cannot find an answer. A model that is allowed to reason its way to a policy will, sooner or later, reason its way to one that does not exist.

Three design decisions that would have contained it

The first is labelling. Truell's own first change was to mark AI responses as AI responses. Users reading the reply as a human staff answer treated it as authoritative, and the whole cascade depended on that reading. A visible label changes how people weigh the answer and gives them a reason to ask a human before cancelling.

The second is a retrieval boundary. If the bot can only state policies that appear in a maintained policy corpus, then an unexplained logout produces either a citation or an admission of ignorance. The admission is far better than the invented rule. It is also cheap to enforce, because you can require the agent to quote the source it relied on and refuse to send a reply that asserts a policy with no quote attached.

The third is escalation on novelty. Multiple users asking the same new question in a short window is a signal that something changed on the product side, and that is exactly when a first-filter bot is least trustworthy. Routing a spike in near-identical questions to a person, before any reply goes out, would have put the race condition in front of an engineer while the number of cancellations was still zero.

What we would want to see measured

Companies running AI on the support inbox tend to report deflection rate, which is the share of tickets closed without a human. This incident shows why that number is dangerous on its own. A bot that confidently closes a ticket with a fabricated answer scores perfectly on deflection and negatively on everything that matters. The metric we would want alongside it is a sampled audit of closed tickets where a person checks whether each factual claim about the product is traceable to a source, and how often the bot said it did not know.

We would also want the cost side stated in the same breath. Cursor's bug affected a narrow set of users on slow connections. The bot's answer affected everyone who read it. If you are going to put a model in front of customers, the question to ask before launch is what the worst plausible wrong answer costs, and whether the system has any way to notice it is giving one.

Sources

  1. Hacker News, Cursor IDE support hallucinates lockout policy, causes user cancellations (with cofounder reply)
  2. Wikipedia, Cursor (code editor), April 2025 support bot incident