Kimi K3 underperforms US frontier AI models in planning cyberattacks

A joint evaluation report has stated that Moonshot AI’s Kimi K3 is not up to par with the strongest American AI systems when it comes to building software exploits and running simulated network attacks.
The United Kingdom’s AI Security Institute (AISI) and the US Center of AI Standards and Innovation (CAISI) ran the evaluation, and they made their findings public on July 23.
Moonshot is expected to open the model’s full weights to the public on July 27, and the findings say that Kimi K3’s safeguards did nothing to stop it from helping with offensive cyber work.
What did the joint evaluation discover about Kimi K3?
The two institutes ran Kimi K3 through ExploitBench, a Carnegie Mellon test that scores how far a model can push an exploit through to completion. It is based on 41 vulnerabilities found in Chrome’s V8 engine after 2023. Kimi K3 scored 32%. The leading models from the US averaged 76.2%, and China’s GLM-5.2 came in at 24%.
Arbitrary code execution, which hands an attacker full control of a target machine, is the benchmark’s most dangerous outcome. US models reached it on 20 of the 41 tasks on average. Kimi K3 reached it on none.
… Continue reading the full article at the original source below.



