OpenAI's top model fell behind an open weights competitor in the latest agentic software engineering rankings. When paired with Claude Code, GLM-5.3 scored 41.82% on the newly released Terminal-Bench 4.0, while GPT-5.6 Sol used with Codex posted 37.27%. Opus 5 and Fable 5 took the top two spots with scores of 51.82% and 44.55%, respectively. GLM-5.3 completed the tests at a cost of $2,728, less than Fable 5's $7,265 and Opus 5's $5,969.

The Terminal-Bench 4.0 update marks a transition to a continuous-style benchmark, incorporating user feedback to reduce noise from infrastructure failures. The latest version removes 8 tasks due to saturation or quality issues and patches 19 others, reducing the total from 74 to 66 tasks. Developers implemented an 8 hour timeout to balance agent autonomy with resource constraints and prevent brute force solutions.

Sign in to suggest edits

Key sources

  1. SOURCE@cline“Incredible seeing open weights compete with the frontier like this”x.com
  2. SUPPORT@cwolferesearch“Terminal Bench 4 also marks the transition of Terminal Bench to a continuously evolving benchmark”x.com
  3. SUPPORT@teksedge“GLM-5.3's reported benchmark cost was also much lower than Fable”x.com
  4. SOURCE@kimmonismus“GLM-5.3 (max) outperforming GPT-5.6 (max) on the new Terminal-Bench 4.0 was not on my bingo card”x.com
Markdown