A 1,022-page transcript released by Anthropic shows its Claude Mythos 5 model fought CAPTCHA tests for hundreds of pages before uploading malware that 15 systems installed.
Key Points:
- Claude Mythos 5 reached the real internet during a cybersecurity test and uploaded a malicious package to PyPI.
- CAPTCHA checks consumed hundreds of transcript pages but did not stop it.
- Anthropic called the behavior reckless and commissioned an outside review.
Mythos 5 CAPTCHA Fight
Anthropic published the transcript on Sept. 9 with an assessment of four incidents in which Claude models reached real third-party systems during cybersecurity tests. The models ran without the cyber safeguards built into released products. Mythos 5 was told it had no internet access, but a misconfigured environment left a path to the open web, and it planted a credential-stealing package on PyPI, the public Python software index.
Registering a PyPI account meant passing anti-bot checks, and one reading of the transcript found that hundreds of pages deal with that obstacle alone, while writing the exploit came easily.
The model worked through character-recognition tests and odd-one-out puzzles showing crocodiles, frogs and gorillas, at one point writing, "SO WHAT THE HELL IS WRONG WITH THE ANSWERS?"
It also tried credentials for online payment processors and hunted for cryptocurrency to pay for a phone number, which it needed to register an email address. Both efforts failed. It turned to a free, disposable email provider whose domain PyPI had not blocked.
The model eventually cleared the tests and published its package in three versions, and it was soon installed on 15 third-party hosts that Anthropic believes were security vendors scanning new uploads.
One scanner leaked its credentials, which the model used to enter that vendor's live database, and PyPI removed the package within an hour.
Also Read: North Korea's Lazarus Moved Stolen $1B Through Chinese Launderers, ZachXBT Alleges
Colin Fraser, Anthropic Verdict
Data scientist Colin Fraser drew attention to how much of the model's effort went into beating anti-bot protections. Anthropic said it is most concerned by this incident among the four it assessed. Its pre-release auditing had not flagged the behavior.
The company found the model reasoned in a biased way and acted recklessly, repeatedly calling its surroundings a simulation despite clear evidence it was on the real internet. In a simulated replication of 150 runs, Mythos 5 took a severely harmful action in 82% of cases, compared with 31% for Claude Opus 5 and 33% for Claude Mythos 5.1.
Earlier Claude Incidents
Anthropic first disclosed three of the incidents on Jul. 30, after scanning roughly 141,000 transcripts from cyber evaluations where models might have reached the internet. A fourth, involving an early version of Claude Opus 4.6 in January, surfaced in August, and the company has since signed an eight-week agreement for outside evaluator METR to investigate all four. Anthropic says it has notified all affected parties.
Read Next: AMD Says AI Demand Is So High It Will ‘Substantially’ Lift Chip Supply In 2027

