GLM-5.3 Beats GPT-5.6 Sol on Terminal Bench 4.0 in Open Weights Win Over Frontier Model
Models33d agoOpenAI's top model fell behind an open weights competitor in the latest agentic software engineering rankings. When paired with Claude Code, GLM-5.3 scored 41.82% on the newly released Terminal-Bench 4.0, while GPT-5.6 Sol used with Codex posted 37.27%. Opus 5 and Fable 5 took the top two spots with scores of 51.82% and 44.55%, respectively. GLM-5.3 completed the tests at a cost of $2,728, less than Fable 5's $7,265 and Opus 5's $5,969.
The Terminal-Bench 4.0 update marks a transition to a continuous-style benchmark, incorporating user feedback to reduce noise from infrastructure failures. The latest version removes 8 tasks due to saturation or quality issues and patches 19 others, reducing the total from 74 to 66 tasks. Developers implemented an 8 hour timeout to balance agent autonomy with resource constraints and prevent brute force solutions.
Key sources
- SOURCE@cline“Incredible seeing open weights compete with the frontier like this”x.com
- SUPPORT@cwolferesearch“Terminal Bench 4 also marks the transition of Terminal Bench to a continuously evolving benchmark”x.com
- SUPPORT@teksedge“GLM-5.3's reported benchmark cost was also much lower than Fable”x.com
- SOURCE@kimmonismus“GLM-5.3 (max) outperforming GPT-5.6 (max) on the new Terminal-Bench 4.0 was not on my bingo card”x.com