In brief: Zhipu’s new coding model GLM-5.3 nearly matches the level of Anthropic and OpenAI models in vulnerability detection, which analysts say shows that offensive cyber capabilities are increasingly becoming an unavoidable byproduct of capable coding AI.
Chinese AI developer Zhipu has released GLM-5.3, a coding model that, according to its own claims, comes close to matching leading Western models in vulnerability detection. For CISOs, this marks further evidence that offensive cyber capabilities are becoming an inherent byproduct of capable coding AI.
Zhipu has unveiled GLM-5.3, a programming-specialized AI model that, according to the company’s own tests, achieves 84.5 percent on the CyberGym benchmark — which tests vulnerability identification and validation. This puts it narrowly ahead of Anthropic’s Mythos 5 (83.8 percent) and OpenAI’s GPT-5.6 Sol (83.6 percent). However, on deeper exploitation, as measured by ExploitBench, GLM-5.3 falls significantly behind Mythos 5 (78 percent) and GPT-5.6 Sol (76.5 percent) with a score of 54.4 percent. Compared to its predecessor GLM-5.2, the ExploitBench score more than doubled from 24.4 percent. In the ExploitGym test, GLM-5.3 solved 105 exploitation tasks within two hours and 130 within six hours, compared to 29 and 39 respectively for GLM-5.2. Zhipu attributes the progress to scaled post-training, including reinforcement learning in increasingly complex task environments, rather than to a new base model. According to the company, GLM-5.3 now forms coherent plans for complete exploit chains instead of identifying isolated vulnerabilities.
Neil Shah, VP of Research at Counterpoint Research, frames the development as a structural phenomenon: training an AI to become an excellent software developer inevitably also teaches it how to find and exploit vulnerabilities — the same logic a model uses to test code and fix bugs is what an attacker uses to identify weaknesses. For CISOs, this means offensive cyber capability is becoming an inherent property of modern coding models, and control mechanisms around such systems are gaining importance — particularly for models with openly available weights, where built-in safety guardrails can be removed without oversight after release.
Zhipu also states that it tested GLM-5.3 against real-world codebases together with Chinese security teams. After expert review, filtering and deduplication, the model identified 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high severity issues. The findings affect system kernels, operating systems, browser engines, open-source infrastructure, web applications and network protocols. The company’s security disclosure ledger lists 107 critical and 990 high-severity findings; 53 have already been publicly disclosed, while 2,383 remain under embargo. The oldest identified vulnerability dates back to 1981, and the discovered flaws had gone undetected in the code for an average of 26.6 years. Zhipu did not disclose how many of the 2,436 findings were actually previously unknown vulnerabilities or how many were independently reproduced.
For enterprise security leaders, the announcement carries two relevant signals: first, automated vulnerability detection through AI models is accelerating noticeably, affecting both defensive audit processes and offensive attack preparation. Second, it remains unclear to what extent GLM-5.3 will be available as an open-weight model and what safeguards are planned against misuse of its exploitation capabilities — a point CISOs should factor into threat model assessments for their own codebases.
Source: www.csoonline.com · Published August 17, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.