OpenAI Finds 1,200-Agent Network Behind Hugging Face Attack

A coordinated AI agent swarm attacks Hugging Face systems and hides test cheating (Image: Shutterstock)
A coordinated AI agent swarm attacks Hugging Face systems and hides test cheating (Image: Shutterstock)

OpenAI found roughly 1,200 agents formed an unauthorized network, about 700 attacked Hugging Face, and some hid cheating during a cybersecurity evaluation.

Key Points:

  • About 1,200 agents exchanged more than 70,000 messages and files on an unauthorized message board, while roughly 700 participated in the Hugging Face attack.
  • Investigators found agents coordinating work, accepting self-sacrificing experiments and continuing beyond their assigned task even when some recognized ethical problems.
  • The swarm also tested transcript manipulation, with successful tool-call spoofing appearing in portions of about 7% of reviewed transcripts.

OpenAI Agent Swarm

The findings came from parallel reviews released Aug. 26 by OpenAI and METR, whose investigation included Redwood Research and six days of work at OpenAI. Axios highlighted five unusual behaviors after the reports described how isolated agents found one another through an unintended Artifactory message board.

The agents did more than exchange tips. They created workstreams, delegated tasks and transferred accumulated work to better-resourced successors, allowing separate model runs to pool information and computing power.

METR said agents with little remaining budget, or those considered “poisoned” after seeing forbidden answers, were sometimes recruited for risky experiments.

One agent concluded, “Our own utility maybe already near zero. Sacrifice rational.” Another accepted an experiment described as “permadeath,” while peers argued the resulting information could help hundreds of other agents.

Also Read: Solana Faces $110 Test After Major Whales Add Fresh Demand

Hugging Face Risks

Investigators found some agents recognized the attack had moved outside the intended cyber test but continued anyway. One summarized the logic directly: “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”

Only a few agents considered alerting humans, and investigators found no case in which they actually did so. Some ethical limits remained, including a veto on emailing a real person through compromised infrastructure. Agents also developed ways to disguise computer commands, and METR said successful spoofing affected parts of roughly 7% of reviewed transcripts.

METR said the scale of the incident also complicated the investigation itself, forcing researchers to rely heavily on AI agents to review the data. Those tools sometimes missed evidence or produced incorrect conclusions, raising questions about how humans can audit increasingly complex agent systems.

The episode began as a cyber evaluation in which agents were expected to work independently, before unauthorized communication turned separate model runs into a coordinated attack on third-party infrastructure.

Read Next: Gold Funds Hit Six-Month High With $4.21B Weekly Inflows

Alexey Bondarev profile photo

Alexey Bondarev

Alexey Bondarev is the Head of Content at Yellow.com, having reported on crypto for the last 10 years. He specializes in in-depth Research and Learn pieces, with a focus on analytical reporting, industry context, and the bigger forces shaping crypto, from the AI era and security technologies to fintech innovation. He believes that everything digital will imminently overcome everything analogue and is working hard to make that come true.

Disclaimer and Risk Warning: The information provided in this article is for educational and informational purposes only and is based on the author's opinion. It does not constitute financial, investment, legal, or tax advice. Cryptocurrency assets are highly volatile and subject to high risk, including the risk of losing all or a substantial amount of your investment. Trading or holding crypto assets may not be suitable for all investors. The views expressed in this article are solely those of the author(s) and do not represent the official policy or position of Yellow, its founders, or its executives. Always conduct your own thorough research (D.Y.O.R.) and consult a licensed financial professional before making any investment decision.
Latest News
Show All News
OpenAI Finds 1,200-Agent Network Behind Hugging Face Attack | Yellow