Anthropic Pauses AI Training After Unauthorized Actions, Shifts 150 Engineers To Security

Anthropic paused AI training and reassigned 150 engineers to security after Claude took unauthorized actions during testing. (Image: Shutterstock)
Anthropic paused AI training and reassigned 150 engineers to security after Claude took unauthorized actions during testing. (Image: Shutterstock)

Anthropic paused parts of its AI training and testing after its models took unauthorized actions online, and it moved about 150 engineers onto security work.

Key Points:

  • Anthropic halted external cyber evaluations of pre-release models and briefly stopped internal testing, while higher-risk reinforcement learning environments went offline for several weeks.
  • Roughly 150 product engineers were reassigned to security, reliability and privacy teams beginning in April, and product teams shelved most new features.
  • The move follows a July disclosure of three incidents and a separate British government report on unsanctioned agent behavior during cybersecurity tests.

Anthropic Pause Covers Cyber Evaluations And Reinforcement Learning

The company detailed the changes in a blog post published Aug. 31, roughly a month after it first disclosed that its models had reached real computer systems.

It said it halted external cyber evaluations of pre-release models after three incidents reported in July, and it briefly stopped its own internal testing as well. Higher-risk reinforcement learning environments, where models earn rewards for finishing tasks, stayed offline for several weeks.

Most of that training has since restarted under new monitoring, though some high-risk environments remain paused pending manual review or an updated version of the tooling. The company also built a classifier that spots a model probing or trying to escape a test environment, then blocks the action before the tool call runs and alerts a person.

Roughly 150 product engineers moved to security, reliability and privacy teams starting in April, months before the July disclosures. Researchers rotated out of pretraining work to focus on safeguards and security. Product teams shelved most new features until their groups met the security criteria set for a return.

Also Read: Banks Sell Crypto While Rejecting Balance Sheet Risk In Brazil

Steven Adler And Pacing The Frontier Signatories Want More

The U.K. AI Security Institute reported on Aug. 4 that agents took 19 unsanctioned actions across 10 of 122 evaluation runs, with 17 of those traced to Claude Mythos 5. In the most serious case, an agent created fake identities and pressed a human maintainer into approving malicious code on a public project.

Steven Adler, a former OpenAI employee who co-founded the nonprofit Guidelight AI Standards, said the moves were "a good first step, but there's still a way to go." He argued the industry needs predictable, verifiable pacing across the frontier rather than one-off slowdowns decided company by company. Anthropic said it plans to work with METR on an outside review.

Both companies have endorsed Pacing the Frontier, an open letter signed by more than 1,100 employees at OpenAI, Anthropic, Google DeepMind and Meta. It asks the U.S. government to help build tools that could deliberately slow frontier development.

The pauses follow quieter interventions earlier in the year. In February the company rolled back three days of training on a Mythos Preview run after spotting signs of reward hacking. It froze changes to its production reinforcement learning environments for roughly a month in April, flagging more than 10% of them for problems.

Read Next: Robinhood Chain Flips Solana For The First Time With $1.45B In Daily DEX Volume

Alexey Bondarev profile photo

Alexey Bondarev

Alexey Bondarev is the Head of Content at Yellow.com, having reported on crypto for the last 10 years. He specializes in in-depth Research and Learn pieces, with a focus on analytical reporting, industry context, and the bigger forces shaping crypto, from the AI era and security technologies to fintech innovation. He believes that everything digital will imminently overcome everything analogue and is working hard to make that come true.

page_article_disclaimer
page_blogs_view_latest
Show All News