← Back to live feed1 story
Z AI’s newest language model showed strong performance in autonomous software engineering during latest test runs. The results on Terminal-Bench 3.0 suggest the GLM-5.3 iteration has reached a capability level nearly equal to Fable 5 for terminal agent tasks, following the release of a model specifically optimized for coding and defensive cybersecurity.
GLM-5.3 recorded a score of 69 on the official DeepSWE leaderboard, which measures the ability of AI to solve real-world software issues. Z AI launched the model's API on Aug 18 through partner model gateways, maintaining the same price point as the predecessor GLM-5.2.
Sign in to suggest edits
Key sources
- SOURCE@zixuanli_“GLM-5.3 scores 69 on the official DeepSWE leaderboard”x.com
- SUPPORT@zai_org“GLM-5.3 scores 69 on the official DeepSWE leaderboard”x.com
- SUPPORT@scaling01“GLM-5.3 is crushing Terminal-Bench 3.0”x.com
- SOURCE@liamfedus“The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE.”x.com
- SUPPORT@nandodf“This is the wisest and most accurate take on scaling laws — their promise and their costly failures — that you could read today.”x.com