← Back to live feed1 story
The GLM-5.3 AI model earned a 64.3 coding score to become the top open-weight system for programming tasks. It currently ranks 13th out of 50 models overall on the Vals Index, climbing from the 18th position held by GLM-5.2, with a total score of 57.0. The model secured the top open-weight spot for legal research and code migration and placed second for the Finance Agent v2 benchmark.
Performance data shows a 95.4% score on SWE-bench and 71.5% on Terminal-Bench 2.1. Vibe Code Bench results rose 14 points to 78.1%. A performance gap remains between GLM-5.3 and Kimi K3 on difficult Terminal-Bench tasks, where GLM-5.3 scored 53.3% compared to Kimi K3's 74.4%.
Sign in to suggest edits
Key sources
- SOURCE@valsai“#1 on our proprietary Legal Research, #1 on Code Migration, and #2 on Finance Agent v2”x.com
- SUPPORT@wesroth“78.1% Vibe Code Bench, up 14 points”x.com
- SUPPORT@zixuanli_“strong capabilities in legal and financial reasoning”x.com
- SOURCE@arena“At a $0.12 median cost per task and a +4.6% net improvement, it has reshaped the Pareto frontier!”x.com
- SUPPORT@dabit3“rolled out support for GLM 5.3 Flash along with an improved model picker UX in Devin CLI”x.com
- SUPPORT@factoryai“GLM-5.3 and GLM-5.3 Flash have arrived in Droid”x.com
- SUPPORT@factoryai“stronger results on complex programming and long-horizon tasks”x.com
- SUPPORT@deeplearningai“This unexpected jump in exploit generation prompted a temporary safety hold on the model weights’ release”x.com