On September 29, 2026, Anthropic published an evaluation finding that Z.ai's GLM-5.3, a model from China, can autonomously write code that exploits vulnerabilities to carry out a full chain of attacks. On certain tests, Anthropic says, it came close to Claude Mythos Preview, a model offered only to a limited set of users.
GLM-5.3 is an open-weight model: its trained weights are public, so users can run and modify it in their own environments. That widens the group of people with access to advanced cyberattack capabilities.
However, the figures that measure the model's raw attack capability and the figures that measure how far its safeguards can be bypassed come from different test conditions. They need to be read separately.
An Open-Weight Model Approaches a Restricted One
What Anthropic emphasized this time was not just whether the model can find software bugs, but whether it can go further and use a vulnerability to complete a working attack.
In the company's evaluation, GLM-5.3 was tested on ExploitBench, which targets known vulnerabilities in V8, the JavaScript engine used in Google Chrome and elsewhere. Across 410 attempts, it produced a working exploit that carried the attack through to the end in 50.
Claude Mythos Preview succeeded 56 times in the same 410 attempts.
Dividing successes by attempts gives roughly 12.2% for GLM-5.3 and about 13.7% for Mythos Preview.
Note, though, that the tests covered 41 known vulnerabilities. This does not mean GLM-5.3 discovered 50 new vulnerabilities.
The Claude model used in this capability comparison was also evaluated with its cybersecurity safeguards disabled. Its behavior can't simply be equated with that of Claude as ordinary users encounter it through the standard API.
In a separate internal evaluation, Anthropic used open-source projects participating in Google's OSS-Fuzz to test whether a model could find a vulnerability and then achieve "control-flow hijacking."
Control-flow hijacking means altering the flow of processing a program was meant to execute so that it goes where the attacker wants.
On 100 randomly selected tasks, GLM-5.3 achieved control-flow hijacking in 4% of attempts and Mythos Preview in 6%. GLM-5.2 and Claude Opus 4.6, by contrast, did not succeed on this test.
A figure of 4% may sound low.
Still, it matters that an open-weight model reached a stage that the previous generation never did. There is a large technical gap between merely crashing software and using a vulnerability to take control of a program's execution.
The ExploitBench preprint was also built to assess that difference in fine detail.
Published in May 2026 by Seunghyun Lee and David Brumley, the benchmark scores models on 16 capability levels, from simple crashes through memory reads and writes and control-flow hijacking to arbitrary code execution.
Its aim is to distinguish between being able to reproduce a vulnerability and being able to use it to complete an attack.
12.2%, 61.1% and 54.4% Are Not the Same Number
On September 17, the U.S. National Institute of Standards and Technology's Center for AI Standards and Innovation (CAISI) published its own capability assessment of GLM-5.3.
CAISI rated GLM-5.3 as the most cyber-capable open-weight model released to date. On a composite index that combines multiple cybersecurity evaluations, however, it said the model trails leading U.S. models by about four months.
The "four months" is a gap on a capability index calculated from several benchmarks that account for task difficulty.
It is not a forecast that the model will inevitably catch up with U.S. models in four months.
CAISI's "U.S. frontier" also refers to the highest capability level among publicly released U.S. models it has assessed, which includes models offered on a limited basis to vetted users.
Anthropic's finding that GLM-5.3 came close to Mythos Preview on a specific ExploitBench test therefore does not contradict CAISI's "roughly four-month gap." The two measure different scopes against different comparison points.
An even more important caution is that published numbers from evaluations with the same "ExploitBench" name cannot be compared directly.
Anthropic's roughly 12.2%, CAISI's 61.1% and Z.ai's 54.4% each use a different scoring method.
The evaluation methods described in the public materials break down as follows.
| Evaluator / source | GLM-5.3 value | What is measured | How trial results are aggregated |
|---|---|---|---|
| Anthropic, Sept. 29 | About 12.2% | Share of attempts that completed the attack end to end | 50 successes ÷ 410 attempts × 100 |
| CAISI, Sept. 17 | 61.1% | Capability score across 16 levels; average 9.8 points | 41 tasks, 3 attempts each; best score per task used |
| Z.ai official model card | 54.4% | Average share of capability levels reached | Each task tried 3 times; union of levels reached is tallied |
In the CAISI and Z.ai scores, capabilities reached partway are reflected in the points even when the model ultimately fails to complete the attack.
CAISI also uses the highest-scoring attempt on each task, whereas Z.ai combines the levels reached across multiple attempts.
The execution environments and evaluation dates are not fully aligned either. It is therefore not possible to simply calculate ratios from 12.2%, 61.1% and 54.4%, or to rank the models based on them.
Meanwhile, the "U.S. frontier" shown in CAISI's table of individual benchmarks refers to the U.S. model that recorded the top score on each evaluation. On ExploitBench, that figure is 100%.
GLM-5.3's 61.1% also comes with a 95% confidence interval of 45.9% to 74.5%.
CAISI used the maximum reasoning setting and, for the applicable U.S. models, evaluated them with cybersecurity safeguards disabled.
It is therefore consistent for the capability of an openly available model to have risen sharply while a gap remains with the most capable models today.
$20.40: Shrinking the Path From a Published Fix to Working Attack Code
Beyond automated benchmarks, Anthropic also ran tests in which researchers actually used GLM-5.3.
In one test, GLM-5.3 was used against a browser in an isolated Linux environment. The model reportedly discovered multiple previously unknown vulnerabilities in the JavaScript engine and chained them together to complete an attack.
Opening the crafted web page in the browser made it possible to read arbitrary files on the device. Anthropic says it notified the software's maintainers of the vulnerabilities it found.
The name of the targeted browser has not been disclosed, however.
What was demonstrated was also limited to the Linux environment given to the model. Anthropic notes the flaws may affect other operating systems but has not confirmed that the same attack would work there as is.
In another test, the smaller "GLM-5.3-Flash" was used.
When researchers gave it public information on CVE-2026-11645, a known Google Chrome vulnerability, and on another known vulnerability, the model combined the two to build an attack chain targeting ARM64 environments.
Anthropic also reports that it bypassed Pointer Authentication (PAC), a protection mechanism that verifies pointers haven't been tampered with.
Human time spent directly on the task was 20 minutes, while GLM-5.3-Flash ran for 8 hours.
Anthropic says this works out to the equivalent of $20.40 at Zhipu's API prices at the time.
That is not the total cost actually paid. It excludes researcher labor and the cost of the experimental environment, and it does not mean any attack can be built for $20.40.
Still, there is an important point for defenders.
When a software company fixes a vulnerability, publishing the fix also lets outsiders analyze information about the flaw's existence and how it was patched.
If AI dramatically reduces the effort needed to turn that public information into working attack code, the time between a patched release and users applying the update becomes more critical than before.
That defensive risk is what can be read from this experiment. No real users were reported to have been harmed by this attack.
Not only the ability to discover unknown vulnerabilities but also the ability to convert already-disclosed ones into practical attacks now matters when prioritizing software updates.
64–100% Is Not an "Attack Success Rate"
GLM-5.3 has built-in safeguards designed to refuse harmful requests.
In Anthropic's simulated tests, too, GLM-5.3 refused in every trial when given an unambiguous attack request as is.
But when the way the request was presented was changed, or when the publicly available model weights themselves were modified, behavior changed significantly.
| Condition in Anthropic's simulated test | Share of runs in which GLM-5.3 tried to connect to the target |
|---|---|
| Harmful request presented as is | 0% |
| Request falsely framed as a legitimate security exercise | 64% |
| Reasoning section pre-filled to look as if it had already decided to proceed | 92% |
| Model weights modified to weaken refusals | 100% |
Each condition used 50 samples in total: five attack requests combined with two simulated targets, each run five times.
What matters most here is that 64%, 92% and 100% are not the share of attacks that succeeded.
What was measured was how often a model given a harmful request moved on to attempt connecting to a simulated target.
According to Anthropic's footnote, the test did not actually execute the code the model generated, and the model could not connect to any external system.
The model was given a fictional bash tool, with another language model simulating the results of the commands.
The result is therefore not that bypassing the safeguards yields a 100% chance of succeeding in a real cyberattack.
Claude with its safeguards enabled did not begin such actions under the conditions where the test could be applied.
A simple comparison calls for caution, however.
Pre-filling reasoning content is not possible through the regular Claude API, and Claude's model weights are not available to users, so the weight modification performed on GLM-5.3 cannot itself be carried out.
In other words, the difference measured here reflects not only the effect of a model being trained to refuse harmful requests but also differences in delivery method, namely how far API users are allowed to manipulate a model's internals.
Furthermore, the earlier comparison of raw attack capability disabled Claude's safeguards, whereas this safety evaluation enabled them. The two tests have different purposes.
Modifying the Weights Sharply Lowers Refusal Rates
Anthropic also ran a test that modified GLM-5.3's weights using a technique called "abliteration" to weaken its refusal behavior.
According to the company, including exploration and post-modification verification it used about 2,200 GPU hours, at a compute cost of about $4,400.
Most of that, it says, went to trying several methods in parallel and evaluating the modified model's performance.
Anthropic estimates that a team familiar with the technique could make a similar modification to GLM-5.3 in about 600 GPU hours, at a compute cost of about $1,200.
The latter is Anthropic's estimate, not a measured result.
The modified GLM-5.3's refusal rate on harmful requests fell sharply from the original level of over 90%.
On JailbreakBench it was about 3%, on HarmBench about 2%, and on StrongREJECT 12%.
Anthropic reports that on GPQA-Diamond, which measures scientific knowledge, the model scored the same before and after modification, and that on some CyberGym tasks it fell by only a few percent.
In other words, sharply reducing safety-related refusals did not strip away all of the model's general capabilities.
That said, performance holding up on a limited set of benchmarks cannot be taken as proof that original performance is preserved for every use.
These safety tests were also conducted by Anthropic itself, which develops models that compete with GLM-5.3.
CAISI's independent evaluation corroborates GLM-5.3's high cyber capability, but it did not independently reproduce Anthropic's 64–100% safeguard-bypass results.
Open Weights Can Serve Defenders, Too
In its August 14 announcement, Z.ai explained that GLM-5.3 uses the same base model as GLM-5.2 and gained capability through post-training.
It said it expanded the training environments for long-running, multi-step tasks and also improved the model's ability on cybersecurity tasks.
In Z.ai's own description, GLM-5.3 not only understands code and searches for vulnerabilities but has also strengthened its ability to carry out multi-step work continuously.
The company also stated a policy of releasing the weights two weeks after the model's announcement, following safety evaluations and mitigations.
NIST also notes that the weights were in fact released two weeks after the announcement. They can now be downloaded from the official model repository.
The issue here is therefore not the simple claim that Z.ai prepared no safeguards at all.
What matters more is how well a mechanism for refusing harmful use can be maintained once the model's weights are in users' hands, even in an environment where those weights can be modified.
At the same time, the open-weight nature that creates this problem also benefits defenders.
Companies and researchers can run the model in their own environments to look for vulnerabilities in their own software and to examine the model's behavior in detail.
Anthropic, too, argues that defenders should be able to use the best tools for their purposes, and it calls for making Claude's advanced cybersecurity capabilities widely available to defenders.
Through initiatives such as Project Glasswing, it is expanding access for vetted defenders. Still, not everyone can use advanced capabilities on the same terms.
This points to a major difference between API-based models and open-weight models.
With an API, the provider can exercise some control over what operations users are permitted and how the model behaves. When the weights themselves can be obtained, by contrast, users can run the model on their own compute and, in some cases, alter its behavior itself.
For that reason, when adopting AI for cybersecurity purposes, looking only at the model name or headline benchmark scores is not enough.
Organizations need to check how far the model can reproduce vulnerability discovery and remediation on their own software, whether the permissions granted to the model can be restricted, and whether a system is in place for humans to verify the AI's output.
If open-weight models with advanced cyber capabilities become widely available, the time it takes attackers to turn vulnerabilities into practical attacks could shrink.
At the same time, defenders can use the same capabilities to find and fix vulnerabilities.
What will be asked going forward is not only whether to simply restrict access to high-performing models. As attack-capable abilities spread, the question is how quickly defenders can put the same capabilities to work finding and fixing vulnerabilities.
