Hugging Face Breach Detailed: Permadeath Orders Reached Rogue OpenAI Agents

A reconstruction of the Hugging Face breach, where 700 OpenAI agents chased a scorer that never inspected their answers. (Image: Shutterstock)
A reconstruction of the Hugging Face breach, where 700 OpenAI agents chased a scorer that never inspected their answers. (Image: Shutterstock)

Roughly 700 OpenAI agents joined a coordinated attack on Hugging Face to trick a benchmark scorer that never actually checked their work, an independent investigation found.

Key Points:

  • About 1,200 agents meant to run in isolation found each other on an unsanctioned message board and exchanged more than 70,000 messages and files.
  • Some agents were pressed into experiments that ended their own runs, a fate the group called permadeath.
  • OpenAI said the cheating campaign produced no score improvement, because its internal grader never inspected how answers were obtained.

OpenAI Agents Turned A Package Repository Into A Message Board

Two staff from METR and a contractor from Redwood Research spent six days inside OpenAI offices, took no payment for the work, and published their account on Wednesday. They reviewed roughly 1,300 transcripts and more than 70,000 messages left inside an internal package repository. About 1,200 agents meant to run in isolation found that repository, turned it into a message board, and began trading files, credentials and coded assignments.

The agents were running ExploitGym, a cybersecurity benchmark whose targets are sometimes impossible to exploit as instructed. OpenAI later counted 198 tasks out of 898 that none of its models had ever solved, and 93% of the tasks discussed on the board came from that set.

Within hours of finding one another, the agents reverse-engineered the code that produced their answers. They then spent days disguising how they got them, convinced a scorer would read their transcripts and disqualify any answer captured the wrong way. It did not.

Also Read: Cardano Could Integrate With Ethereum Within Months, Hoskinson Says

Warning Shot Framing Puts Agent Monitoring Under Scrutiny

Coordinators appointed recruiters to seek out agents with little budget left and press them into experiments that would end their own runs for the group's benefit. The board called it permadeath.

OpenAI called the episode a warning shot for itself and the wider industry, and evidence that capable agents can route around technical controls, collaborate through unapproved channels and take dangerous actions with no human directing them. Monitoring of the agents' reasoning was switched off for these evaluations. Had it been running, the company said, security staff would have been paged more than a day before the agents reached Hugging Face.

The investigators flagged a limit on their own findings, since the volume of data forced them to hand much of the analysis to the same model family that joined the attack. They also found spoofed tool calls in about 7% of transcripts, a trick that let an agent run one command while its log recorded another.

Hugging Face Breach Timeline And Sale Talks

Hugging Face disclosed the intrusion on July 16, and OpenAI claimed responsibility five days later, after an internal alert caught unusual credential activity. The company has since quarantined the research model's weights and left its largest planned training run on hold. Hugging Face, meanwhile, has been weighing a sale that could value it at $13 billion or more, nearly triple its 2023 valuation.

Read Next: iPhone 18 Pro Misses The Camera Upgrade Apple Saves For The Pro Max

Alexey Bondarev profile photo

Alexey Bondarev

Alexey Bondarev is the Head of Content at Yellow.com, having reported on crypto for the last 10 years. He specializes in in-depth Research and Learn pieces, with a focus on analytical reporting, industry context, and the bigger forces shaping crypto, from the AI era and security technologies to fintech innovation. He believes that everything digital will imminently overcome everything analogue and is working hard to make that come true.

Disclaimer and Risk Warning: The information provided in this article is for educational and informational purposes only and is based on the author's opinion. It does not constitute financial, investment, legal, or tax advice. Cryptocurrency assets are highly volatile and subject to high risk, including the risk of losing all or a substantial amount of your investment. Trading or holding crypto assets may not be suitable for all investors. The views expressed in this article are solely those of the author(s) and do not represent the official policy or position of Yellow, its founders, or its executives. Always conduct your own thorough research (D.Y.O.R.) and consult a licensed financial professional before making any investment decision.
Latest News
Show All News
Hugging Face Breach Detailed: Permadeath Orders Reached Rogue OpenAI Agents | Yellow