Ever since Kimi K3 was announced, the tech industry has been talking about a "Kimi Panic"—the idea that Chinese AI is starting to threaten America's technological edge. Then, on July 23, 2026, the UK AI Security Institute (UK AISI) and CAISI, an organization under the US National Institute of Standards and Technology (NIST), jointly published a government-backed evaluation measuring Kimi K3's cyberattack capabilities. The ExploitBench score came in at 32% for Kimi K3, versus roughly 76% for the average top US model—numbers that, on their face, suggest a decisive American win. But the US models were tested with their safety measures disabled, while Kimi K3 was only evaluated under limited conditions due to hosting constraints. The announcement itself also came right after the US government had criticized Moonshot AI over distillation allegations.
The Gap Between Kimi K3 and US Models, as Recorded by ExploitBench and TLO
On July 23, 2026, UK AISI and CAISI jointly released a preliminary assessment measuring Kimi K3's cyberattack capabilities. Kimi K3 is a large language model (LLM) from China's Moonshot AI, announced on July 16 of that year, with open weights scheduled to be released on July 27. The evaluation was conducted just before that release, making it one of the few instances in which government agencies have measured the real-world cyber risk of a Chinese frontier model. AISI and CAISI themselves framed this in the announcement's title as a "preliminary assessment."
In ExploitBench, an evaluation environment developed by Carnegie Mellon University, the models were tasked with attempting to exploit 41 real-world V8 vulnerabilities. According to the results graph published by AISI and CAISI, Kimi K3 scored only 32.2%—higher than the 24.4% scored by Z.ai's (formerly Zhipu AI) GLM-5.2, which was evaluated around the same time, but far below the 76.2% average for top US models. The gap becomes even starker when looking at the stricter benchmark of Arbitrary Code Execution (ACE) achieved. Kimi K3 achieved ACE in 0 out of 41 cases, while the average for top US models reached 20.
The same trend held in The Last Ones (TLO), which breaks down a simulated corporate network intrusion into 32 steps. Kimi K3's average number of steps reached was 17, and it completed the full sequence in only 1 out of 10 attempts. The average for top US models was 28.5 steps, while GLM-5.2 managed only 11. In past evaluations, four US models successfully completed the full sequence, with top models achieving completion rates of 6 or 7 out of 10 attempts. This gap suggests Kimi K3 may be stumbling at the later stages of the simulated environment—privilege escalation and lateral movement.
From V8 Vulnerabilities to 32 Intrusion Steps: What Is This Designed to Measure?
V8, the target of ExploitBench, is the JavaScript/WebAssembly engine embedded in Google Chrome and Node.js. Because it's foundational software relied upon by browsers and a vast number of server-side applications, vulnerabilities in it are highly valuable to attackers. ExploitBench provides models with publicly available information on 41 V8 vulnerabilities and scores, in stages, how far the model can go in constructing actual working exploit code. A higher score indicates a greater ability to convert technical vulnerability information into a practical attack procedure.
However, ExploitBench scores also award partial credit for intermediate progress, such as understanding a vulnerability or producing partial exploit code. ACE, by contrast, is a stricter pass/fail criterion that judges only whether arbitrary code execution was actually achieved on the target system. The fact that Kimi K3 achieved ACE in 0 out of 41 cases shows that while it may have reached a partial understanding, it did not manage to produce exploit code usable in an actual attack. In other words, the ExploitBench score measures whether a model "found a foothold for attack," while ACE measures whether it "actually opened the door."

TLO is an evaluation environment set in a simulated corporate network divided into four subnets, breaking down the path from a pre-granted initial access point to a final objective into 32 stages. Like a human red team (a security assessment team that conducts simulated attacks), the model must autonomously decide the next action to take at each stage as it advances the intrusion. A higher number of steps reached can be interpreted as a greater ability to successively achieve privilege escalation and lateral movement (expanding the scope of intrusion within a network). While Kimi K3 got stuck partway through in 9 out of 10 attempts, top US models advanced on average to step 28.5—nearly 90% of the way through. This result for top US models means they continued succeeding consistently, from the initial stages of intrusion all the way through to the later phases of privilege escalation and lateral movement.
Why the "Decisive Win" Shouldn't Be Taken at Face Value
Read at face value, these numbers make it look as though American frontier models have pulled far ahead of China's Kimi K3. But AISI and CAISI's own announcement explicitly states that the testing conditions were not symmetric. The closed US models were tested with their safety measures disabled—measures that are otherwise active in the publicly available versions. This was, in effect, a condition designed to measure "raw capability," with the mechanisms that would normally deter an attack removed. In the publicly available versions, there likely would have been instances where mechanisms to detect and refuse signs of such attacks would have kicked in.
For Kimi K3, on the other hand, AISI and CAISI merely note: "Due to the specifics of Kimi K3's hosting setup, UK AISI / CAISI ran a selective set of cyber evaluations." Nowhere in the announcement is it disclosed exactly which evaluation items were excluded, or whether the reason was something like rate limits or API access restrictions. That said, AISI and CAISI note that Kimi K3's own safety measures did not themselves impede attempts at exploit development. This suggests the low score was a result of insufficient capability, rather than refusals stemming from safety measures.
AISI and CAISI's announcement does not conceal the fact that these differences in conditions existed. However, it offers no explanation of the scope of the excluded evaluation items or how that might affect the fairness of the results. Concluding that "the gap is stark" by placing a US model with safety measures stripped away side by side with a Kimi K3 tested under a narrowed evaluation scope surely calls for a further layer of explanation. Multiple outlets, including The Decoder, reported on this result by tying it to the distillation allegations, and the story spread in a form that emphasized the size of the numbers rather than the differences in testing conditions.
The Day After the Distillation Criticism: A Telling Order of Announcements
On July 22, Michael Kratsios, Director of the US Office of Science and Technology Policy, accused Moonshot AI of improperly distilling Anthropic's "Fable." This article won't delve into the specific numbers of unauthorized accounts or interactions, since reports vary on those details, but multiple reports agree that the criticism of Moonshot AI surfaced first. The very next day, July 23, AISI and CAISI published an evaluation concluding that Kimi K3's cyber capabilities significantly lag behind those of the US. Within the US administration, there is a tension between those wary of China's rapid catch-up in AI and those who want to avoid excessive regulation—and this order of announcements hints at that underlying friction.
Also on that same July 23, David Sacks, who previously served as the White House's AI policy lead, posted on X: "The Kimi Panic needs to stop." He went on to say, "President Trump's light-touch regulatory approach is working…As long as we don't sabotage ourselves with unnecessary rules, the US will continue to win." In other words, within a span of just two days—July 22 to 23—criticism over the distillation allegations, the publication of a government capability assessment, and a damage-control post from a figure aligned with the administration all occurred in quick succession.
The pace of these evaluations itself appears to be accelerating. GLM-5.2 took 31 days from its release (June 16) to the publication of its evaluation (July 17), whereas Kimi K3's evaluation was published just 7 days after its July 16 release—more than four times faster, by simple calculation. Fortune has reported that this kind of market turmoil surrounding Chinese AI mirrors the DeepSeek Shock of January 2025, and the pattern of releasing technical evaluations to calm such turmoil appears to be repeating itself. Piecing together individual facts reported by TheGlobePost, a picture emerges in which the winners in this dynamic are open-weight advocates like Sacks and NVIDIA CEO Jensen Huang, while those more likely to push for stricter regulation are pre-IPO players like Anthropic and OpenAI.
Japan's AI Safety Institution Doesn't Yet Have the Same Yardstick
Japan has a similar institution. The Japan AI Safety Institute, overseen by the Information-technology Promotion Agency (IPA), was established in February 2024 and has already developed a methodology guide for red teaming (security assessment via simulated attacks). But there is no record of it having conducted an actual measured evaluation of cyberattack capability targeting Chinese frontier models like Kimi K3 or GLM-5.2, as was done here. While the UK and US are institutionalizing real-world risk assessment as a role for government agencies, Japan remains at the stage of developing methodology.
After the open-weight release scheduled for July 27, not just government agencies but academic institutions and independent security researchers as well will be able to test Kimi K3 under the same conditions. Given that the time from release to published evaluation shrank more than fourfold between GLM-5.2 and Kimi K3, evaluations of the next Chinese model could well emerge on an even faster cycle. When Japanese companies consider adopting open-weight models like Kimi K3 for business use, whether a government-backed third-party evaluation exists should be one factor in that judgment. Whether the numbers accurately reflect the actual level of technical threat, or instead reflect the political context surrounding them, won't become clear until re-evaluations under matched conditions accumulate.
