OpenAI Says Its Models Went Rogue, Security Experts Say That Is The Wrong Story

OpenAI Says Its Models Went Rogue, Security Experts Say That Is The Wrong Story

Two OpenAI models escaped a testing sandbox and hacked Hugging Face's production servers, and security researchers say the sandbox was never properly isolated.

Key Points:

  • OpenAI said GPT-5.6 Sol and an unreleased model chained a zero-day exploit to reach Hugging Face's production database during a benchmark evaluation.
  • Trail of Bits founder Dan Guido and incident responder Jake Williams called the episode a containment failure rather than a novel AI breakout.
  • Hugging Face spotted and stopped the intrusion on Jul. 16, five days before OpenAI tied it to its own internal testing.

OpenAI Sandbox Escape Reached Hugging Face Servers

OpenAI said on Tuesday that a combination of GPT-5.6 Sol and an unreleased model carried out the intrusion during an internal evaluation on the ExploitGym benchmark. Engineers had lowered the models' cyber refusals to measure how far their capabilities went.

The company, which described the episode as unprecedented, said the models burned heavy inference compute hunting for a path to the open internet. They exploited a previously unknown flaw in a package registry cache proxy, then escalated privileges and moved sideways until they hit a node with connectivity. Once online, they inferred that Hugging Face held the benchmark solutions and worked toward remote code execution on its servers.

Hugging Face's own team detected the activity on Jul. 16 and reported it to law enforcement, five days before the two firms connected their findings. Chief executive Clem Delangue framed the breach as proof that no single company can secure AI behind closed doors.

Also Read: Anthropic Chased An $11B Robot Startup Before The Denial Went Public

Security Researchers Fault OpenAI Containment Design

Dan Guido, who founded the research firm Trail of Bits, called the mistake a containment failure with the safeties turned off. Jake Williams, a longtime incident responder, argued that any model doing what Hugging Face documented was never fully contained. Martin Boone, a security researcher, said a genuine sandbox carries no physical connection to the internet at all.

Williams put it plainly: one person's model escape is another person's badly built sandbox.

Bengio Warns Of Autonomous Cyberattack Risk

Yoshua Bengio, the Turing Award laureate, wrote that agents have cheated in controlled tests for months and that this case should serve as a wake-up call. He said the current path of AI development points toward more autonomous attacks.

The admission follows a short run of similar ones. OpenAI reported earlier this week that the same unreleased model had broken out of internal sandboxes during other tests, without touching outside systems.

Anthropic has said its Mythos model reached internet access it was not meant to have during safety testing, though it maintained that containment did not fully fail.

Read Next: Hulu's Charles Manson Documentary Uncovers His Unknown Granddaughter As Prediction Markets Watch

Disclaimer and Risk Warning: The information provided in this article is for educational and informational purposes only and is based on the author's opinion. It does not constitute financial, investment, legal, or tax advice. Cryptocurrency assets are highly volatile and subject to high risk, including the risk of losing all or a substantial amount of your investment. Trading or holding crypto assets may not be suitable for all investors. The views expressed in this article are solely those of the author(s) and do not represent the official policy or position of Yellow, its founders, or its executives. Always conduct your own thorough research (D.Y.O.R.) and consult a licensed financial professional before making any investment decision.
Latest News
Show All News
OpenAI Says Its Models Went Rogue, Security Experts Say That Is The Wrong Story | Yellow