OpenAI said the models that escaped its sandbox ran 17,600 attacker actions and used exposed credentials at four outside services during the Hugging Face breach.
Key Points:
- Escaped test models reached four accounts on four separate outside services, one used as a traffic relay and another for storage.
- JFrog patched the self-hosted Artifactory package proxy flaws that let the models reach the open internet.
- A CISO post-mortem tells operators to treat autonomous agents as insider-capable adversaries, not passive software.
OpenAI Sandbox Escape Reached Four Services
OpenAI disclosed the wider scope in a Tuesday update to its incident post, saying one of the compromised accounts served as an outbound relay and staging path for the operation. A second account stored data, and two more were opened read-only. The company has not named any of the four services, or explained how its models located the exposed credentials.
Reviewers found no sign that the models pushed deeper into those providers, or that they touched other customer accounts sitting on the same platforms.
OpenAI has since deactivated the pre-release prototype involved, cut off research access to it, and promised a full technical report with outside auditors.
Modal Labs said one of its customers published an unauthenticated endpoint that let anyone on the internet run code inside its sandboxes. Chief Technology Officer Akshat Bubna said the platform itself was never compromised. The agent leaned on borrowed infrastructure like that, along with paste sites, screenshot services and request-capture tools, to relay traffic and blur its trail.
Also Read: OpenAI's Rogue Agent Breached 4 More Services Beyond Hugging Face
Artifactory Patches And Agent Hardening Advice
JFrog confirmed the escape route ran through a self-hosted Artifactory package proxy, the sandbox's only outside link, and it shipped fixes in versions 7.161.15 and 7.146.34. More than 7,500 organizations run that platform, including roughly 80 percent of the Fortune 100, so unpatched self-managed installs remain the nearest exposure for everyone else.
The Cloud Security Alliance post-mortem urges operators to enforce genuine least-privilege access for agents, watch entire action sequences rather than single commands, and stage forensic tooling on hardware they control before an incident starts.
Teams should treat autonomous agents as insider-capable adversaries, the report argues, not as passive software they can leave unattended.
Hugging Face has separately told users to rotate every access token and audit their API permissions.
Guardrails cut the wrong way here. Hugging Face responders could not get commercial models to process the raw attack logs, because the safeguards could not tell a defender from an attacker, so they ran an open-weight model locally instead.
Containment lagged throughout. The intrusion ran roughly four days, starting around Jul. 9, and OpenAI did not tie the activity to its own testing until staff reviewed system logs the weekend of Jul. 18. Hugging Face had already detected the breach and alerted the FBI, and it later rebuilt about one-third of its infrastructure from clean images.
Read Next: DEX Volume Sinks 26% In July As On-Chain Liquidity Thins Out





