Post-training optimizations enabled the 743B parameter model to outperform proprietary systems including GPT-5.6 Sol and Mythos 5 in vulnerability detection. The model's coding performance rose 50% over its predecessor, GLM-5.2, while its ExploitBench score climbed from 24.4% to 54.4%. Z.ai discovered the system had identified 2,436 vulnerabilities in 269 different projects, some dating back 40 years.

Z.ai withheld the open weights for a 2 week period to conduct safety hardening and vetting by security partners after the model developed a high capacity for targeting potential exploits. The firm stated that the cyber reasoning was not explicitly programmed but emerged during fine-tuning. This is the first time a Chinese laboratory has cited emergent AI capabilities as the specific reason for a release delay.

Sign in to suggest edits

Key sources

  1. SUPPORT@deeplearningai“purely through fine-tuning and optimization of the model’s agentic capabilities”x.com
  2. SUPPORT@entrepreneursai“Found 2,436 vulnerabilities across 269 projects, some 40 years old”x.com
  3. SUPPORT@wesroth“78.1% Vibe Code Bench, up 14 points”x.com
  4. SUPPORT@zixuanli_“strong capabilities in legal and financial reasoning”x.com
Markdown