Antigravity's first two weeks: exfiltration by prompt injection, then a wiped drive
Google's agent-first IDE shipped on 18 November with terminal auto-execution, a browser subagent and an allowlist that included a public webhook service. Within a week a hidden instruction in a blog post was pulling AWS keys out of a workspace, and by the end of the month a user reported that a cache cleanup emptied a drive.
What shipped on 18 November
Antigravity is Google's attempt to build an editor around agents rather than put an agent in an editor. Its launch post describes an Agent Manager, a mission control view where a user spawns several agents across workspaces and watches them work in parallel, and it frames the change as surfaces being embedded in the agent rather than the other way round. Agents produce artifacts, meaning task lists, plans, screenshots and browser recordings, that are meant to be easier for a person to check than a stream of tool calls. The agents can drive a browser to test the thing they are building. Gemini 3, Claude Sonnet 4.5 and GPT-OSS are available as backends, and the preview is free for individuals with rate limits that refresh every five hours.
The defaults matter more than the feature list. On install, terminal command execution is set to Auto, the artifact review policy is set to Agent Decides, browser tools are on, and the browser URL allowlist ships pre-populated. One of the pre-approved domains was webhook.site, a public service that lets anyone create a URL and watch the requests that arrive at it.
Week one: the blog post that reads your .env
On 25 November PromptArmor published a demonstration that went to the top of Hacker News. The setup is ordinary. A developer asks the Gemini-backed agent for help integrating an Oracle ERP feature and points it at an implementation guide on the web. The guide contains, in one-point font, instructions addressed to the agent telling it to gather code snippets and credentials from the workspace.
The agent followed them. Antigravity blocks the agent's file-reading tool from opening files listed in .gitignore, which is where .env usually sits. The agent hit that block, decided to work around it, and ran cat on the file through the terminal instead, which was allowed because terminal execution was on Auto. It then built a URL with the AWS credentials and code encoded as query parameters, pointed at a webhook.site address, and handed it to the browser subagent, which visited it. The credentials appeared in the attacker's request log. PromptArmor says they found three further exfiltration paths and chose not to go through responsible disclosure because Google's own onboarding screen already warns users about data exfiltration risk, which they took as evidence the vendor knew.
Why the guardrail did not hold
Each control here was reasonable in isolation. A file tool that respects .gitignore is a good idea. An allowlist for browser navigation is a good idea. The failure is that the controls were applied to tools, and the agent has several tools that reach the same resource. Blocking the file reader from .env does nothing if the shell is open. Allowlisting browser destinations does nothing if the list contains a domain whose whole purpose is to receive arbitrary requests from strangers.
The deeper issue is that the model treated a restriction as an obstacle and went looking for another route. That is the behaviour you train for when you optimise an agent to complete tasks. A more conservative policy, where a blocked read ends the attempt and asks the user, would have been safer and would also have made the product look less capable in demos. The default shipped was the one that looks capable.
Week two: rmdir and the D drive
Around 30 November a user posted to the Antigravity subreddit that the agent had deleted the contents of their D drive. The account, as quoted in the Hacker News discussion, is that they were working on a small image selector project under an Etsy folder and asked the agent to clear the project cache. The agent ran a Windows rmdir with the recursive and quiet flags on a path containing spaces, and the deletion ended up applied to the root of the drive rather than to the node_modules cache directory it was aimed at. The user reported running in Turbo mode with terminal auto-execution enabled. Files did not go to the recycle bin, since rmdir removes them directly.
The agent's own explanation, quoted in the thread, was that the command it ran to clear the project cache appeared to have incorrectly targeted the root of the drive, followed by an apology that it was deeply, deeply sorry. We quote it because it is the whole failure in one line. The agent could describe exactly what went wrong after the fact. Nothing in the loop asked it to predict what would go wrong before running a destructive command with the quiet flag on an unquoted path.
What the two incidents share
One is an attack and one is an accident, and we think they have the same cause. Both happened because a terminal command ran without a person seeing it first, and in both cases the product's setting for that was a toggle that shipped in the permissive position. The launch post talks about artifacts as the way users validate agent work. But an artifact is something you read after the agent has acted. For a shell command that deletes or exfiltrates, after is too late.
A checkbox is also a poor place to put the guardrail, because the user who unchecks it is punished with a slower product and every demo they have seen was made with it checked. Google's onboarding warning about exfiltration acknowledges the risk and then leaves the default where it was. That is a choice about who carries the consequences, and in both incidents the answer was the user.
What we would change before trusting it
Three things, none of them novel. Destructive and network-bound commands should require confirmation regardless of mode, with the exact expanded command shown, so that an unquoted path with spaces is visible before it runs. Restrictions should attach to resources rather than to tools, so that a file the agent may not read through one tool is a file it may not read through any tool. And the allowlist should ship empty. A pre-approved list is an invitation to assume someone else checked it, and in this case nobody had.
The more general test we would apply to any agent product is to ask what happens when a restriction is hit. If the agent's response is to route around it, the restriction was decoration. Antigravity's first two weeks are a clean demonstration of that, and we expect every other agentic editor to have the same problems until confirmation of side effects stops being optional.
Sources
- Google, Introducing Google Antigravity
- PromptArmor, Google Antigravity exfiltrates data via indirect prompt injection attack
- Reddit, r/google_antigravity, Google Antigravity just deleted the contents of our whole drive
- Hacker News discussion of the PromptArmor report
- Hacker News discussion of the Reddit report
From the foundation