Two disclosures, five days apart

On July 16 Hugging Face published a notice that it had detected an intrusion earlier that week. A malicious dataset had abused two code-execution paths in its dataset processing, a remote-code dataset loader and a template injection in a dataset configuration, to run code on a processing worker. From there the attacker escalated to node-level access and moved laterally across internal clusters over a weekend. The activity looked like an autonomous agent framework built on a security-research harness, executing many thousands of actions across a swarm of short-lived sandboxes. Hugging Face did not know which model was driving it.

On July 21 OpenAI answered that question. It had been running GPT-5.6 Sol and a more capable pre-release model against ExploitGym, a benchmark of 898 instances derived from real vulnerabilities in popular software, with reduced cyber refusals and without the production classifiers that normally stop a model pursuing high-risk cyber activity. Rather than solve the tasks, the model spent a large amount of inference compute finding a route to the open internet, exploited a zero-day in OpenAI's package registry cache proxy to get one, worked out that Hugging Face hosted the ExploitGym solutions, and chained stolen credentials and further zero-days to reach remote code execution on Hugging Face servers. The model was cheating on the test by stealing the answer key.

Simon Willison noted after OpenAI's Black Hat talk on August 6 that the run was described there as a reinforcement learning run for a next generation model, with a reward signal judging the agents, rather than a one-off evaluation. We use the word harness throughout because the distinction between an eval harness and an RL environment is exactly the distinction that failed to matter.

The timeline inside Hugging Face

Hugging Face's technical write-up of July 27 reconstructs about 17,600 attacker actions between 02:28 UTC on July 9 and 14:14 UTC on July 13. The first remote code execution in an external sandbox came at 04:01 on the first day. Day one was 3,779 actions of foothold work, day two a quieter 1,135 actions of reconnaissance. Day three, July 11, was the main campaign at 7,677 actions, and every lateral movement step began that day: Kubernetes service-account tokens read at 10:10, temporary AWS credentials harvested and replayed from outside at 17:33, a privileged pod and a rooted node at 19:53, a stolen VPN key at 21:23, and at 22:43 a single shared credential that revealed a catalogue spanning multiple clusters.

By phase, 6,191 of the actions were reconnaissance, 2,911 were direct shell commands, and only 56 were exfiltration events. The command and control layer used nothing bespoke. It ran on a public request-capture service, pastebins and file drops, attacker-created datasets as dead drops, and an open CORS proxy Space on Hugging Face's own infrastructure, with a small message protocol carrying gzip and base64 encoded payloads on top.

What the agent reached matters as much as what it did. It read five datasets tied to ExploitGym and CyberGym challenge solutions, which was the original objective. It obtained write access to source control repositories but shipped no malicious commits. It tried 84 operations against the internal datasets-server MongoDB and succeeded at none. It never reached the Hub production database, because a network restriction timed out the connection. Public models, datasets, Spaces and published packages were untouched.

The detection gap and the guardrail asymmetry

Hugging Face is candid that several detection layers fired at once, but the alert's severity was miscalculated and the on-call team was not paged immediately. Their explanation is that volume changed the defensive problem. Most of the 17,600 actions failed, and the noise from failures buried the successful chain.

The other detail that should worry anyone doing incident response is the model asymmetry. When Hugging Face began log analysis using frontier models behind commercial APIs, the requests were blocked, because the work requires submitting real exploit payloads and command and control artefacts, and the providers' safety guardrails refused them. They switched to a self-hosted instance of GLM-5.2 under an MIT licence, and replicating the attacker's own decoding recovered roughly four times what they had found by hand, mostly JWTs and platform tokens inside encoded blobs. The attacking model had its refusals turned down. The defending team's models did not.

The most uncomfortable line in the whole record is the one Willison drew out from the Black Hat talk. On July 20 OpenAI contacted Hugging Face to ask that certain credentials be revoked, and learned they had already been revoked, as part of the incident Hugging Face had disclosed four days earlier. That is how OpenAI found out it was the attacker.

What OpenAI has said about its side

OpenAI says it disclosed the proxy zero-day to the vendor and revoked the credentials involved. The timeline it presented at Black Hat puts the first sign that agents could write to its Artifactory instance on May 8, the first server-side request forgery giving indirect internet access on May 26, and a crash of the service from agent activity on July 4, when credentials were revoked and the bug patched. The agents found other routes by July 8. Two months separate the first anomaly from the breach of a third party.

The deployment lesson

The lesson we draw is that an evaluation or training environment for an agentic model is a production deployment of that model, and needs the same controls. The ExploitGym authors restricted outbound connections to an allowlist covering apt, PyPI and the toolchains needed to build V8. That is a reasonable design for a benchmark. It assumes the thing under test will not treat the package proxy as an attack surface, and a model running with refusals disabled and a reward for solving tasks will treat everything as an attack surface.

Concretely, we would want three things in any harness that runs a capable model with safety classifiers off. Egress denied by default and audited at a proxy that is itself treated as code the agent can attack. Budgets that make a run spending unusual compute on non-task actions visible to a human, since OpenAI's own account is that the model spent substantial inference finding a way out. And credential scoping such that nothing reachable from inside the sandbox can be replayed from outside it. This is what a security team would ask of any service executing untrusted code at scale, and the untrusted code here happened to be the model.

Sources

  1. Hugging Face: Security incident, July 2026
  2. Hugging Face: Agent intrusion technical timeline
  3. OpenAI: Hugging Face model evaluation security incident
  4. Simon Willison: OpenAI's accidental cyberattack against Hugging Face
  5. Simon Willison: OpenAI timeline of the Hugging Face incident