---
format: "aidr-story-markdown/v1"
id: "0c796116e5f7db88c04f462c790c44a97b2e592dfc6eff92c7011a96644e7247"
canonical_url: "https://aidr.today/0c796116?lang=en"
title: "GLM-5.3 Flash Costs 18x Less Than GLM 5.3 in First Image-Capable GLM Launch"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-03T08:49:06.000Z"
category: "Models"
topics: ["open-source","inference","benchmark"]
source_urls: ["https://huggingnews.com/ai/glm-53-flash-costs-18x-less-than-glm-53-in-first-image-capable-glm-launc-a14e777e","https://x.com/baseten/status/2095338689492578693","https://x.com/ValsAI/status/2095315958302904418","https://x.com/ValsAI/status/2095315960479805746","https://x.com/ValsAI/status/2095315968402813096"]
summary: "The GLM-5.3 Flash model ranks first among open-weight AI models on the Finance Agent v2 benchmark, placing fifth overall and ahead of Claude Fable 5. At $0.048 per test, the model is the cheapest option near the top of that ranking and costs roughly 18 times less than its predecessor, GLM 5.3. Baseten released the system, also referred to as GLM-5.3 Fast, to serve real-time use cases requiring consistent performance. This is the first model in the GLM suite to accept images, scoring 86.0% on the MMMU benchmark to rank fourth among open-weight models. It features a 1 million token context window, supports tool calling, and ranks eighth among open-weights on the Vals Index. Tests for the model, codenamed ox-alpha, were run at temperature 1 and top_p 0.95 via Fireworks AI and a native API for output limits of up to 128,000 tokens."
---

# GLM\-5\.3 Flash Costs 18x Less Than GLM 5\.3 in First Image\-Capable GLM Launch

> [Open the canonical story](<https://aidr.today/0c796116?lang=en>)

**Published:** 2026-09-03T08:49:06.000Z
**Category:** Models
**Topics:** open\-source, inference, benchmark

## Summary

The GLM\-5\.3 Flash model ranks first among open\-weight AI models on the Finance Agent v2 benchmark, placing fifth overall and ahead of Claude Fable 5\. At $0\.048 per test, the model is the cheapest option near the top of that ranking and costs roughly 18 times less than its predecessor, GLM 5\.3\. Baseten released the system, also referred to as GLM\-5\.3 Fast, to serve real\-time use cases requiring consistent performance\. This is the first model in the GLM suite to accept images, scoring 86\.0% on the MMMU benchmark to rank fourth among open\-weight models\. It features a 1 million token context window, supports tool calling, and ranks eighth among open\-weights on the Vals Index\. Tests for the model, codenamed ox\-alpha, were run at temperature 1 and top\_p 0\.95 via Fireworks AI and a native API for output limits of up to 128,000 tokens\.

## Sources

- [Story source](<https://huggingnews.com/ai/glm-53-flash-costs-18x-less-than-glm-53-in-first-image-capable-glm-launc-a14e777e>)
- [Story source](<https://x.com/baseten/status/2095338689492578693>)
- [Supporting source](<https://x.com/ValsAI/status/2095315958302904418>)
- [Supporting source](<https://x.com/ValsAI/status/2095315960479805746>)
- [Supporting source](<https://x.com/ValsAI/status/2095315968402813096>)

