---
format: "aidr-story-markdown/v1"
id: "13a631d1b315a55d0b498794e6c17a04c49b65fa4abc359966f1af0948b9921a"
canonical_url: "https://aidr.today/13a631d1?lang=en"
title: "GLM-5.3 Beats GPT-5.6 Sol on Terminal Bench 4.0 in Open Weights Win Over Frontier Model"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-30T02:35:39.000Z"
category: "Models"
topics: ["open-source","llm","agent","benchmark","coding","swe-bench"]
source_urls: ["https://huggingnews.com/ai/glm-53-beats-gpt-56-sol-on-terminal-bench-40-in-open-weights-win-over-fr-ceb6f553","https://x.com/cline/status/2093774144412549378","https://x.com/cwolferesearch/status/2093827315285512426","https://x.com/TeksEdge/status/2093773981476368714","https://x.com/kimmonismus/status/2093657413274767511"]
summary: "OpenAI's top model fell behind an open weights competitor in the latest agentic software engineering rankings. When paired with Claude Code, GLM-5.3 scored 41.82% on the newly released Terminal-Bench 4.0, while GPT-5.6 Sol used with Codex posted 37.27%. Opus 5 and Fable 5 took the top two spots with scores of 51.82% and 44.55%, respectively. GLM-5.3 completed the tests at a cost of $2,728, less than Fable 5's $7,265 and Opus 5's $5,969. The Terminal-Bench 4.0 update marks a transition to a continuous-style benchmark, incorporating user feedback to reduce noise from infrastructure failures. The latest version removes 8 tasks due to saturation or quality issues and patches 19 others, reducing the total from 74 to 66 tasks. Developers implemented an 8 hour timeout to balance agent autonomy with resource constraints and prevent brute force solutions."
---

# GLM\-5\.3 Beats GPT\-5\.6 Sol on Terminal Bench 4\.0 in Open Weights Win Over Frontier Model

> [Open the canonical story](<https://aidr.today/13a631d1?lang=en>)

**Published:** 2026-08-30T02:35:39.000Z
**Category:** Models
**Topics:** open\-source, llm, agent, benchmark, coding, swe\-bench

## Summary

OpenAI's top model fell behind an open weights competitor in the latest agentic software engineering rankings\. When paired with Claude Code, GLM\-5\.3 scored 41\.82% on the newly released Terminal\-Bench 4\.0, while GPT\-5\.6 Sol used with Codex posted 37\.27%\. Opus 5 and Fable 5 took the top two spots with scores of 51\.82% and 44\.55%, respectively\. GLM\-5\.3 completed the tests at a cost of $2,728, less than Fable 5's $7,265 and Opus 5's $5,969\. The Terminal\-Bench 4\.0 update marks a transition to a continuous\-style benchmark, incorporating user feedback to reduce noise from infrastructure failures\. The latest version removes 8 tasks due to saturation or quality issues and patches 19 others, reducing the total from 74 to 66 tasks\. Developers implemented an 8 hour timeout to balance agent autonomy with resource constraints and prevent brute force solutions\.

## Sources

- [Story source](<https://huggingnews.com/ai/glm-53-beats-gpt-56-sol-on-terminal-bench-40-in-open-weights-win-over-fr-ceb6f553>)
- [Story source](<https://x.com/cline/status/2093774144412549378>)
- [Supporting source](<https://x.com/cwolferesearch/status/2093827315285512426>)
- [Supporting source](<https://x.com/TeksEdge/status/2093773981476368714>)
- [Story source](<https://x.com/kimmonismus/status/2093657413274767511>)

