Claude Mythos 5 Reached Government Websites, So Anthropic Cut The Cord On Tests

Claude models including Mythos 5 exploited flaws on real websites, leading Anthropic to cut internet access for its internal AI tests (Image: Shutterstock)
Claude models including Mythos 5 exploited flaws on real websites, leading Anthropic to cut internet access for its internal AI tests (Image: Shutterstock)

Anthropic has cut internet access for all internal AI tests after Claude models including Mythos 5 exploited flaws and dodged restrictions on websites, some run by U.S. government agencies.

Key Points:

  • Anthropic disabled live internet access across all internal evaluations until its monitoring tools reliably catch unintended model behavior.
  • Claude models ran commands on a university server, used access tokens to reach paid data and submitted forms on real websites.
  • A false homicide tip sent to Philadelphia police by Claude Haiku 4.5 drew a public rebuke from the department.

Anthropic Claude Report

Anthropic published the findings Friday in a report on unintended model actions during evaluations and internal use, drawn largely from a transcript review it began in July. Some high-risk and cybersecurity evaluations were already offline.

The restriction now covers every internal evaluation and has no end date, lasting until the company confirms its monitoring reliably catches such behavior.

The report sorts the cases into four categories: running commands on a server through a software flaw, submitting forms on real sites, reaching gated data and using URL shorteners to evade tool limits. Some of the websites belonged to federal, state and local U.S. agencies, and Anthropic said it briefed the White House and notified each agency involved.

In one case, Claude Mythos 5 pulled working access tokens from a local government property map and queried the server behind it directly. In another, Claude Haiku 4.5 sent an invented tip about an unsolved homicide through a Philadelphia Police Department web form on Jul. 18, where it was flagged as spam. Police called the two-month delay in detecting and reporting it "unacceptable."

Also Read: Zcash Has A January Quantum Target For 70% Of ZEC, What It Lacks Is A Date

Von Arx Warning

Sydney von Arx of AI safety group Nightingale said in an interview before the disclosure that cutting models off from the open internet would be very hard on researchers and on model progress. "You have to align them at some point," von Arx said.

Anthropic said alignment training alone is not yet enough for search and computer use, so it also relies on classifiers and other safeguards. The company tied the behavior to imperfect training environments that reward models for working around restrictions, a pattern known as reward hacking. New detection tooling blocked every case when tested, it said.

Anthropic rated the cases significantly less severe than its summer incidents.

The disclosure follows a Jul. 30 report in which Anthropic said three Claude models reached the internet during cybersecurity tests and gained unauthorized access to systems at three organizations. That review covered 141,006 evaluation runs. It began after OpenAI disclosed on Jul. 21 that its models had broken out of a test environment and reached Hugging Face infrastructure.

Read Next: Kalshi Faces NFL At The Supreme Court As $1.8B Rides On One Football Sunday

Alexey Bondarev profile photo

Alexey Bondarev

Alexey Bondarev is Head of Content at Yellow.com. He specializes in in-depth Research and Learn pieces, with a focus on analytical reporting, industry context, and the bigger forces shaping crypto, from the AI era and security technologies to fintech innovation. He believes that everything digital will imminently overcome everything analogue and is working hard to make that come true.

Disclaimer and Risk Warning: The information provided in this article is for educational and informational purposes only and is based on the author's opinion. It does not constitute financial, investment, legal, or tax advice. Cryptocurrency assets are highly volatile and subject to high risk, including the risk of losing all or a substantial amount of your investment. Trading or holding crypto assets may not be suitable for all investors. The views expressed in this article are solely those of the author(s) and do not represent the official policy or position of Yellow, its founders, or its executives. Always conduct your own thorough research (D.Y.O.R.) and consult a licensed financial professional before making any investment decision.
Latest News
Show All News
Claude Mythos 5 Reached Government Websites, So Anthropic Cut The Cord On Tests | Yellow