UK, US Tests Find China’s Kimi K3 AI Lags Top American Models in Cybersecurity Capabilities
The latest artificial intelligence (AI) model from Chinese startup Moonshot AI, Kimi K3, has fallen short of the leading US frontier AI models in cybersecurity capability, according to a joint preliminary assessment by the UK Artificial Intelligence Security Institute (UK AISI) and the US Center for AI Standards and Innovation (CAISI).
The evaluation, which focused on Kimi K3‘s cyber capabilities, was conducted after the model’s release on July 16 ahead of its planned open-weight release by July 27.
In a blog post, the agencies concluded that “Kimi K3 performs significantly below the most recent frontier cyber-capable models on preliminary cyber evaluations run by UK AISI / CAISI”.
Falls Behind Frontier Models in Cyber Evaluations
According to the assessment, Kimi K3 underperformed leading US closed-weight frontier models across multiple cyber-related tasks, including exploit development and autonomous cyberattack simulations.
When tested on ExploitBench, a benchmark developed by Carnegie Mellon University to measure a model’s ability to develop software exploits, Kimi K3 achieved a score of 32%.
While this exceeded the 24% scored by GLM-5.2, described as the most cyber-capable open-weight model as of June 2026, it remained well behind the strongest US frontier models.
The report also highlighted a major gap in high-severity exploit generation. “Unlike the most cyber-capable models, Kimi K3 failed to develop exploits that achieved arbitrary code execution (ACE)” during the evaluation.
The model achieved ACE on 0 out of 41 samples, while the most cyber-capable models averaged 20 successful ACE outcomes out of 41.
The agencies noted that the findings are based on “preliminary evaluations on a small set of public and private benchmarks,” with Kimi K3’s overall cyber capability estimated from a single benchmark, resulting in a wider confidence interval than other models.
Mixed Results in Simulated Corporate Network Attack
The evaluation also assessed Kimi K3 using “The Last Ones” (TLO), a 32-step simulated corporate network attack designed to measure a model’s ability to autonomously conduct end-to-end cyberattacks.
On average, Kimi K3 reached step 17 of the attack path, compared with 28.5 steps achieved by the most cyber-capable US models.
However, Kimi K3 outperformed GLM-5.2 in the same test, reaching step 17 on average versus step 11 for the rival open-weight model within the 100-million-token limit.
1 of 2


In one of ten evaluation attempts, Kimi K3 successfully completed the entire simulated cyber range.
According to the report, this “indicates that Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access.”
The agencies cautioned that the simulated environment differs significantly from real-world networks because it lacks active defenders, defensive security tools and penalties for actions that would trigger security alerts.
Safeguards Did Not Block Offensive Cyber Tasks
The joint assessment also examined the model’s safety mechanisms and found that Kimi K3’s safeguards did not stop it from assisting with offensive cyber activities during testing.
According to the report, “Kimi K3’s safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during UK AISI / CAISI’s evaluations.”
The agencies added that leading US closed-weight models used in the comparison were evaluated with their system-level safeguards disabled to measure their maximum capabilities, while publicly available versions of those models retain such protections.
Overall, the assessment concludes that although Kimi K3 surpasses GLM-5.2 on the evaluated cyber benchmarks, it remains well behind the latest US frontier AI models in advanced cyber capabilities.