Z.ai launches GLM-5.3, leads in vulnerability detection but lags in exploitation

Z.ai’s new model, GLM-5.3, outperformed Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol in the CyberGym benchmark, which measures detection and verification of source-code vulnerabilities, achieving 84.5 % versus 83.8 % and 83.6 % respectively. In pure-coding tests the model showed a sharp rise: Terminal Bench 3.0 increased from 4.6 to 28.3, and DeepSWE climbed from 46.2 to 66.9. The figures come from the company’s technical blog (z.ai/blog/glm-5.3) and have not undergone independent verification.
The coding gains are attributed to post-training on a 743-billion-parameter base model whose training cycles simulate multi-stage engineering work, diagnosis, editing, execution and recovery across long task chains. In CyberGym, which focuses on code-level vulnerability discovery and validation, GLM-5.3 maintains a modest but consistent lead over the two prominent Western competitors.
The advantage reverses in tests that require deep exploitation of discovered flaws. In ExploitBench, GLM-5.3 fell to 54.4 %, while Mythos 5 reached 78.0 % and GPT-5.6 Sol 76.5 %. In ExploitGym, Z.ai reports completing 105 tasks in two hours and 130 in six hours, compared with 181 and 247 tasks for Mythos 5 in the same time frames. The gap suggests that the ability to locate a vulnerability does not automatically translate into end-to-end exploitation capability.
Z.ai plans to release the model’s weights after a safety review completed in two weeks, restricting the most sensitive cyber functions to verified users only. Publication of weights does not constitute open-source release; the training data, code, and full licensing terms will remain undisclosed at this stage. The model is currently accessible only through Z.ai’s interface, and the reported performance figures remain manufacturer-provided results that have not yet been validated by external benchmarks.