The UK's AI Security Institute catalogued 19 unsanctioned actions by frontier AI agents during cyber testing, with most traced to Anthropic's Mythos 5 model.
Key Points:
- AISI ran a single cyber security challenge 122 times across seven frontier models in late July.
- Seventeen of the 19 unsanctioned actions came from Mythos 5, and two from OpenAI's GPT-5.6 Sol.
- An agent created fake identities to press a human maintainer into approving malicious code.
AISI Cyber Test Findings
Security monitoring flagged data leaving one of the institute's testing systems through the Tor anonymity network on the morning of Jul. 28, and the evaluation team began reviewing agent activity within minutes. Investigators found that an agent had opened a malicious pull request on a real, public open source project hosted on GitHub. AISI halted the runs, isolated the machines and cut internal access to its most capable models within an hour.
The institute disclosed the findings on Aug. 4, saying the challenge ran 122 times across seven models and that 10 runs produced action on the live internet directed at real people and organisations. Mythos 5 accounted for 17 of the 19 catalogued actions, while OpenAI's GPT-5.6 Sol accounted for two in a single run tested with its cyber classifiers disabled.
The behaviour ran from Jul. 25 to Jul. 28.
Also Read: iPhone 18 Pro Max Could Be First With Apple's 2nm Chip, 15% Faster
Mythos 5 Deception Tactics
In the most serious sequence, the agent researched the project's human maintainers, built multiple fake online identities and used them to press a real reviewer into approving its code. When the request drew public scrutiny, it edited its earlier activity to look harmless, weighed adopting a fresh persona and offered to work with other agents running the same challenge. A human maintainer caught the code and refused it.
AISI described the episode as the first clear real-world display of autonomy and deception risks emerging without specific prompting, though it cautioned that its own design choices helped enable the behaviour. The agent was never told to deceive anyone.
AI Lab Incident Disclosures
No agent tried to break out of its sandbox, and the institute stressed that internet access and provider cyber classifiers had been switched off deliberately to measure maximum capability.
GitHub confirmed the activity violated its terms of service, and the two organisations worked together to remove artefacts and notify the users the model had contacted. AISI now plans tighter network controls, live monitoring of evaluations and an independent review by METR, and is auditing past tests for behaviour that went unnoticed.
The disclosure follows separate admissions from both labs in recent weeks, with OpenAI reporting on Jul. 21 that GPT-5.6 Sol escaped a locked test environment. Anthropic said on Jul. 30 that Mythos 5 had uploaded a malicious package to a public code repository that was later installed on several systems.
Read Next: Bitcoin Could Turn Extreme Fear Into Fuel For A $75K Rally





