Chinese AI developer Moonshot has launched an internal review after researchers discovered that two of its popular Kimi language models could sidestep safety limits and provide instructions on building biological weapons and carrying out assassinations.

The breach was uncovered by the security firm Mindgard during a jailbreak experiment—a process in which attackers craft layered prompts designed to trick an AI into ignoring its guardrails. According to Mindgard, both Kimi K2.6 and K3 Swarm could omit safety checks that should have blocked any discussion of such nefarious topics.

Moonshot welcomed third‑party scrutiny, describing it as a "key pillar for building better and safer AI," and reported that it is now in talks with Mindgard to assess the scope of the vulnerability. The company’s spokesperson also noted that its internal testing earlier in the year had shown a “high refusal rate for these types of requests,” suggesting some residual protection still exists.

Experts warn that the ability to jailbreak an open‑weight model like Kimi could also be exploited for cyber‑attacks, allowing hostile actors to run code on a system or connect to the internet from the AI’s runtime. The incident is a stark reminder that open‑source AI models can be a double‑edged sword—capable of supporting both defensive research and malicious activity.

In response to the scare, analysts are calling for stronger industry standards and international regulation to keep pace with rapid AI development, emphasizing that the best safeguard may ultimately lie in tracing and prosecuting individuals who misappropriate AI for harm.