---
format: "aidr-story-markdown/v1"
id: "36eb2f9145f4b3192fbf02785dc5a187ed3a3d62b92e7cf630c32c1fba93ed6c"
canonical_url: "https://aidr.today/36eb2f91?lang=en"
title: "Z.ai GLM-5.3 Hits 84.5% on CyberGym, First Chinese Lab Delay for Emergent Safety"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-01T15:02:14.000Z"
category: "Models"
topics: ["llm","safety","reasoning","open-source"]
source_urls: ["https://huggingnews.com/ai/update-zai-glm-53-hits-845percent-on-cybergym-first-chinese-lab-delay-fo-9350a61c","https://x.com/DeepLearningAI/status/2094780425319153840","https://x.com/EntrepreneursAI/status/2094576043793301914","https://x.com/WesRoth/status/2094643915290784223","https://x.com/ZixuanLi_/status/2094623238299021469"]
summary: "Post-training optimizations enabled the 743B parameter model to outperform proprietary systems including GPT-5.6 Sol and Mythos 5 in vulnerability detection. The model's coding performance rose 50% over its predecessor, GLM-5.2, while its ExploitBench score climbed from 24.4% to 54.4%. Z.ai discovered the system had identified 2,436 vulnerabilities in 269 different projects, some dating back 40 years. Z.ai withheld the open weights for a 2 week period to conduct safety hardening and vetting by security partners after the model developed a high capacity for targeting potential exploits. The firm stated that the cyber reasoning was not explicitly programmed but emerged during fine-tuning. This is the first time a Chinese laboratory has cited emergent AI capabilities as the specific reason for a release delay."
---

# Z\.ai GLM\-5\.3 Hits 84\.5% on CyberGym, First Chinese Lab Delay for Emergent Safety

> [Open the canonical story](<https://aidr.today/36eb2f91?lang=en>)

**Published:** 2026-09-01T15:02:14.000Z
**Category:** Models
**Topics:** llm, safety, reasoning, open\-source

## Summary

Post\-training optimizations enabled the 743B parameter model to outperform proprietary systems including GPT\-5\.6 Sol and Mythos 5 in vulnerability detection\. The model's coding performance rose 50% over its predecessor, GLM\-5\.2, while its ExploitBench score climbed from 24\.4% to 54\.4%\. Z\.ai discovered the system had identified 2,436 vulnerabilities in 269 different projects, some dating back 40 years\. Z\.ai withheld the open weights for a 2 week period to conduct safety hardening and vetting by security partners after the model developed a high capacity for targeting potential exploits\. The firm stated that the cyber reasoning was not explicitly programmed but emerged during fine\-tuning\. This is the first time a Chinese laboratory has cited emergent AI capabilities as the specific reason for a release delay\.

## Sources

- [Story source](<https://huggingnews.com/ai/update-zai-glm-53-hits-845percent-on-cybergym-first-chinese-lab-delay-fo-9350a61c>)
- [Supporting source](<https://x.com/DeepLearningAI/status/2094780425319153840>)
- [Supporting source](<https://x.com/EntrepreneursAI/status/2094576043793301914>)
- [Supporting source](<https://x.com/WesRoth/status/2094643915290784223>)
- [Supporting source](<https://x.com/ZixuanLi_/status/2094623238299021469>)

