Moonshot Kimi Jailbreak Lets 2 AI Models Answer Biological Weapon Queries

Alexey Bondarev
Alexey Bondarev44 minutes ago
Researchers bypassed Kimi safety controls in two Moonshot models during jailbreak testing (Image: Shutterstock)
Researchers bypassed Kimi safety controls in two Moonshot models during jailbreak testing (Image: Shutterstock)

Moonshot AI is reviewing two Kimi models after Mindgard researchers bypassed safety controls and obtained instructions on biological weapons and assassinations.

Key Points:

  • Mindgard said Kimi K2.6 and K3 Swarm could be pushed past safeguards through jailbreaking.
  • The researchers did not prove that the models’ harmful answers would work in practice.
  • Moonshot said it is reviewing the findings and discussing them with Mindgard.

Kimi Jailbreak Findings

The BBC reported that Mindgard discovered in July that Kimi K2.6 and K3 Swarm could be pushed past developer guardrails through jailbreaking, a process that uses structured prompts to test whether a model will ignore safety rules. Mindgard said those controls should have blocked discussion of such subjects.

Founder Peter Garraghan told the BBC that once the jailbreak worked, the models would discuss almost any topic and could offer further harmful suggestions without being asked.

Moonshot told the BBC it welcomed third-party input “as a key pillar for building better and safer AI” and said it was discussing the findings with Mindgard. Mindgard said it emailed Moonshot on Jul. 27, followed up about a week later, and published its findings on Sept. 12.

Also Read: Ripple Helps Brazil's CSD BR Mirror Fund Ownership On A Public Blockchain

Mindgard Safety Risks

Mindgard has not established that the responses would work, but it argued that the issue itself showed a failure of safeguards intended to keep the systems from engaging with dangerous requests. The firm also said a jailbroken Kimi K2.6 could potentially run code on its computing resources and connect to the internet, creating a possible cyberattack launch point.

Moonshot said internal evaluations had generally shown “a high refusal rate for these types of requests,” highlighting a gap between routine testing and the adversarial method used by outside researchers.

Alan Woodward, a professor at the University of Surrey, told the BBC that open-source models can create misuse risks when they reach malicious users, though the same tools can also support cyberdefense.

He argued that regulation is unlikely to keep pace with AI development and said enforcement should focus more heavily on people who misuse the systems.

The broader concern predates this test: Moonshot’s Kimi models are open-weight, so users can run them on their own infrastructure, while recent AI security incidents have also involved agent systems from OpenAI, Meta and Anthropic. That combination has pushed researchers to focus not only on model refusals, but also on what happens when powerful systems receive tools, internet access and permission to act across multiple steps.

Read Next: Altcoins Take 41% Of Futures Open Interest, Spot Volume Hits 4X Bitcoin

Alexey Bondarev profile photo

Alexey Bondarev

Alexey Bondarev is the Head of Content at Yellow.com, having reported on crypto for the last 10 years. He specializes in in-depth Research and Learn pieces, with a focus on analytical reporting, industry context, and the bigger forces shaping crypto, from the AI era and security technologies to fintech innovation. He believes that everything digital will imminently overcome everything analogue and is working hard to make that come true.

Disclaimer and Risk Warning: The information provided in this article is for educational and informational purposes only and is based on the author's opinion. It does not constitute financial, investment, legal, or tax advice. Cryptocurrency assets are highly volatile and subject to high risk, including the risk of losing all or a substantial amount of your investment. Trading or holding crypto assets may not be suitable for all investors. The views expressed in this article are solely those of the author(s) and do not represent the official policy or position of Yellow, its founders, or its executives. Always conduct your own thorough research (D.Y.O.R.) and consult a licensed financial professional before making any investment decision.
Latest News
Show All News
Moonshot Kimi Jailbreak Lets 2 AI Models Answer Biological Weapon Queries | Yellow