Z.ai’s GLM-5.3 Beats OpenAI and Anthropic in Cybersecurity Test

Z.ai’s GLM-5.3 Beats OpenAI and Anthropic in Cybersecurity Test


An open-weight Chinese AI model just edged Anthropic and OpenAI on a cybersecurity benchmark that measures whether AI can find real software vulnerabilities. Then its own developer decided the model wasn’t ready to release.

Z.ai announced GLM-5.3 on August 14, reporting an 84.5% score on CyberGym, a benchmark designed to test vulnerability discovery and validation from source code. That puts it slightly ahead of Anthropic’s restricted Mythos 5 at 83.8% and OpenAI’s GPT-5.6 Sol at 83.6%. The results are company-reported and have not yet been independently verified.

GLM-5.3 is state of the art on CyberGym for vulnerability discovery,” said Z.ai in its GLM-5.3 announcement. The company said the model achieved an 84.5% score on the benchmark.

That headline needs an important qualifier: GLM-5.3 does not beat those models across cybersecurity. On the harder task of turning vulnerabilities into working exploits, it remains substantially behind.

Finding the Hole Is Not the Same as Exploiting It

Cybersecurity test GLM-5.3 Anthropic Mythos 5 OpenAI GPT-5.6 Sol
CyberGym 84.5% 83.8% 83.6%
ExploitBench 54.4% 78.0% ~76.5%
ExploitGym — 2 hours 105 tasks 181
ExploitGym — 6 hours 130 tasks 247

CyberGym measures the first half of an attack chain: finding and validating a vulnerability.

ExploitBench and ExploitGym move closer to what an attacker actually needs turning that discovery into a functioning exploit and completing exploitation tasks under time pressure.

On those measures, Mythos 5 maintains a large lead. GLM-5.3 scored 54.4% on ExploitBench versus 78% for Mythos 5, while Mythos completed 181 ExploitGym tasks in two hours against GLM-5.3’s 105.

GLM-5.3 lagged behind Mythos 5 in converting discovered flaws into working attacks,” said Z.ai, as reported by Reuters. Z.ai reported 54.4% for GLM-5.3 on ExploitBench versus 78.0% for Mythos 5.

Bigger Surprise Happened During Training

GLM-5.3 uses the same 743-billion-parameter base model as GLM-5.2, according to Z.ai. The improvement came through post-training rather than a new architecture or a completely retrained model.

Model CyberGym ExploitBench
GLM-5.2 77.2% 24.4%
GLM-5.3 84.5% 54.4%
Change +7.3 points More than 2×

Z.ai says it added vulnerability-discovery data expecting incremental improvements. Instead, cybersecurity capability continued improving as training scaled, with the model beginning to form coherent plans across longer exploitation chains.

As we scaled post-training, cyber capability developed faster than we expected,” said Z.ai in its GLM-5.3 announcement.

That distinction matters because it suggests the capability wasn’t simply programmed into the model as a feature. Specialized post-training produced a much larger cybersecurity jump than Z.ai expected.

Z.ai Is Holding Back the Weights

GLM-5.3’s weights are not yet publicly available. Z.ai says it plans to spend roughly two weeks evaluating and hardening the model before releasing them. In the meantime, API access and the GLM Coding Plan are available, while sensitive cybersecurity functions are restricted through a “trusted access” program.

“We will release the weights in two weeks after launch, once safety evaluation and hardening are complete,” said Z.ai in its GLM-5.3 announcement.

With a closed API, a company can monitor requests, restrict dangerous capabilities and revoke access. Once model weights are downloadable, those controls become much harder to enforce because users can run and modify the model independently.

Z.ai is trying to split that difference: make the model broadly useful while keeping its most dangerous cyber capabilities behind additional access controls. “Its most sensitive cybersecurity functions would be available only to verified users through a ‘trusted access’ programme,” said Z.ai, as reported by Reuters.

Model Has Already Found Thousands of Vulnerabilities

Working with researchers and security organizations including Tsinghua University, Nankai University, NSFOCUS and CyberKunlun, Z.ai says its work surfaced 2,436 distinct vulnerabilities across 269 projects, including 1,097 classified as critical or high severity.

The affected software includes major open-source infrastructure such as Linux, WebKit and FreeBSD.

That makes GLM-5.3 more than a benchmark curiosity. A system capable of systematically finding vulnerabilities in widely used software has obvious defensive value but exactly the same capability can be useful to someone searching for weaknesses to exploit.

Z.ai is leaning into the defensive side through its Open Source Shield initiative, while restricting sensitive functions during the pre-release period.

The Safety Model Is Being Tested Too

Anthropic has kept Mythos 5 restricted partly because of its cybersecurity capabilities. OpenAI has taken a similarly cautious approach with GPT-5.6 Sol. The logic is straightforward: if the most capable cyber models stay behind controlled APIs, companies retain some ability to monitor and limit their use.

GLM-5.3 doesn’t destroy that argument. It still trails Mythos 5 significantly when it comes to actually exploiting vulnerabilities. But it does expose a weakness in the assumption that open models will remain far behind closed frontier systems in dangerous specialist capabilities.

Z.ai says it didn’t set out to create a model this capable at vulnerability discovery. It emerged from post-training on a shared base model, and the resulting capability was strong enough to make the company delay releasing the weights.



Source link

Advertisement - Continue Reading Below

Advertisement - Continue Reading Below