← Back to live feed1 story
The GLM-5.3 model from Z AI scored 71.5% on Terminal-Bench 2.1, placing it second among open-weight models behind Kimi K3. The model also earned an 84.8% on LegalBench, the third-highest score for an open-weight model behind Kimi K3 and MiniMax M3, where it performed most effectively on conclusion and issue-spotting tasks.
The Terminal Bench ranking reflects a 19 spot climb from GLM 5.2, with the cost per test falling $0.12 to $0.31. ValsAI noted that current export control restrictions limited evaluations to public and academic benchmarks, and plans to run the Vals Index and Finance Agent benchmarks once the model weights are released.
Sign in to suggest edits
Key sources
- SOURCE@valsai“GLM 5.3 is #2 on Terminal Bench, #3 on Legal Bench, and #6 on Skills Bench among-open weight models”x.com
- SUPPORT@valsai“At just $0.31/test, it is $0.12 cheaper than 5.2”x.com
- SUPPORT@valsai“It is strongest on conclusion and issue-spotting tasks, with interpretation being the weak spot”x.com
- SUPPORT@valsai“We look forward to running our private benchmarks such as the Vals Index, Finance Agent, and Vibe Code Bench once the weights become available”x.com