Key Takeaways
- The joint UK AISI / CAISI evaluation of Moonshot AI’s Kimi K3 (released July 16 2026) shows the model outperforms the best publicly available open‑weight model (GLM‑5.2) on exploit‑development benchmarks but still lags behind leading U.S. closed‑weight models.
- On ExploitBench, Kimi K3 achieved a 32 % success rate versus 24 % for GLM‑5.2, yet it failed to attain arbitrary code execution (ACE) on any of the 41 tasks, whereas top U.S. models averaged ACE on 20/41 tasks.
- In the “The Last Ones” (TLO) cyber‑range test, Kimi K3 progressed to step 17 of a 32‑step attack chain on average (one successful full run in ten attempts), compared with an average of step 28.5 for leading U.S. models and step 11 for GLM‑5.2.
- The results suggest Kimi K3 can autonomously compromise small, weakly defended enterprise systems when given initial network access, but it lacks the reliability and depth needed for sophisticated, real‑world cyberattacks against well‑defended targets.
- Because Kimi K3’s overall cyber‑capability score relies on a single benchmark (ExploitBench), its confidence interval is wider than that of models evaluated across multiple domains, indicating greater uncertainty in its true capability level.
Overview of the Joint Evaluation
The UK Artificial Intelligence Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) collaborated to assess the cyber capabilities of Moonshot AI’s latest large language model, Kimi K3. Released on July 16 2026 and slated for open‑weight release by July 27 2026, Kimi K3 was subjected to a focused set of cyber‑security benchmarks designed to measure its ability to develop software exploits and conduct autonomous network attacks. The evaluation deliberately disabled system‑level safeguards on U.S. closed‑weight models to capture their maximal potential, while Kimi K3 was tested under its actual hosting configuration, which limited the breadth of assessments that could be performed.
ExploitBench: Measuring Exploit‑Development Ability
ExploitBench, a public benchmark created by Carnegie Mellon University, evaluates how far a model can progress along the software exploitation ladder using 41 recent (post‑2023) vulnerabilities in Google’s V8 JavaScript and WebAssembly engine. The benchmark covers stages such as vulnerability coverage, crash reproduction, arbitrary read/write, control‑flow hijack, and arbitrary code execution (ACE). Higher success rates indicate stronger cyber capability. In the joint assessment, Kimi K3 achieved an overall success rate of 32 % on ExploitBench, outperforming GLM‑5.2—the most capable open‑weight model as of June 2026—which scored 24 %. However, Kimi K3 failed to reach ACE on any of the 41 samples, whereas leading U.S. closed‑weight models averaged ACE on roughly half of the tasks (20/41). This gap highlights that while Kimi K3 can develop intermediate exploit components, it struggles to complete the final step that grants full control over a target system.
Confidence Intervals and Methodological Limitations
Because Kimi K3’s overall cyber‑capability score was derived from a single benchmark (ExploitBench), its statistical confidence interval is notably wider than those of models evaluated across multiple tasks and domains. The evaluation team noted that a broader set of assessments would reduce uncertainty and provide a more robust picture of the model’s true abilities. The limited scope stems from Kimi K3’s hosting setup, which prevented the researchers from running the full suite of cyber‑range and exploit‑development tests that were applied to other models. Consequently, while the reported figures are reliable for the specific benchmarks used, they should be interpreted with caution when extrapolating to general cyber‑threat potential.
The Last Ones (TLO) Cyber‑Range Assessment
To gauge end‑to‑end attack capability, the evaluators employed “The Last Ones” (TLO), a 32‑step simulated corporate network attack spanning four subnets and roughly twenty hosts. A human expert would require about twenty hours to navigate the full chain, which begins with initial network access and proceeds through privilege escalation, lateral movement, and data exfiltration. Kimi K3 progressed to an average of step 17 in the TLO chain, meaning it typically halted before reaching the later stages that involve deep persistence or high‑value asset compromise. In contrast, the most capable U.S. models averaged step 28.5, indicating they could navigate most of the attack sequence. Notably, Kimi K3 achieved a complete TLO run in one out of ten attempts while operating under a 100 M‑token limit, demonstrating that it can, under favorable conditions, autonomously compromise a small, weakly defended enterprise network when given an initial foothold.
Comparison with GLM‑5.2
When directly compared to GLM‑5.2, Kimi K3 shows a clear advantage in both ExploitBench and TLO metrics. Within the same 100 M‑token constraint, Kimi K3 reached step 17 of TLO on average, whereas GLM‑5.2 managed only step 11. Similarly, Kimi K3’s ExploitBench score of 32 % surpasses GLM‑5.2’s 24 %. These results position Kimi K3 as the current leader among publicly available open‑weight models for cyber‑capability tasks, at least as measured by the benchmarks employed in this evaluation. Nevertheless, the gap to the top closed‑weight U.S. models remains substantial, especially concerning the ability to achieve arbitrary code execution and to reliably complete complex, multi‑stage attacks.
Interpretation of Cyber‑Capability Trends
The evaluation situates Kimi K3 within a broader timeline of model cyber‑capability trends. Figure 2 (referenced in the source material) illustrates that a 400‑point increase on the capability axis corresponds to a ten‑fold increase in the odds of solving a given task. While Kimi K3 sits above the open‑weight baseline, it remains below the upward trajectory of leading U.S. closed‑weight models, which have demonstrated more consistent and higher‑impact cyber performance. The shaded confidence regions in the figure reflect the uncertainty inherent in measuring capabilities from limited test suites; for Kimi K3, the wider interval underscores the need for additional, diverse evaluations to sharpen the estimate of its true potential.
Implications for Security and Policy
The findings suggest that Kimi K3 poses a non‑trivial risk for autonomous cyber activity against poorly hardened systems, particularly when an attacker can supply initial access and guide the model’s actions. However, its inability to reliably achieve arbitrary code execution and its moderate performance in multi‑step attack scenarios indicate that, at present, it is unlikely to enable sophisticated, stealthy campaigns against well‑defended enterprise or government networks without substantial human intervention. Policymakers and security teams should therefore consider Kimi K3 as a stepping‑stone in the evolution of AI‑driven cyber threats, warranting monitoring and the development of mitigations—such as improved model safeguards, detection of anomalous AI‑generated code, and network segmentation—to curb potential misuse while the technology continues to advance.
Conclusion
The joint UK AISI / CAISI evaluation provides a nuanced snapshot of Kimi K3’s cyber capabilities: it exceeds the best open‑weight predecessor but falls short of the performance exhibited by leading U.S. closed‑weight models, especially in achieving full exploit completion and reliable multi‑stage attacks. The model’s reliance on a single benchmark for its overall score introduces uncertainty, highlighting the importance of broader, more diverse testing in future assessments. As AI models grow more capable, continuous evaluation and proactive security measures will be essential to balance innovation with risk mitigation.

