---
format: "aidr-story-markdown/v1"
id: "f1fbcda63fb8cafbc199ca30eb765a4690b4f25ee816b073cc4352fa68629d40"
canonical_url: "https://aidr.today/f1fbcda6?lang=en"
title: "Z.ai Releases GLM-5.3-Flash After Ox Alpha Processed 20T Tokens in 6 Days"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-26T19:04:30.000Z"
category: null
topics: ["llm","open-source","inference"]
source_urls: ["https://huggingnews.com/ai/zai-releases-glm-53-flash-after-ox-alpha-processed-20t-tokens-in-6-days-262155ed","https://x.com/Zai_org/status/2092616204787626030","https://x.com/ZixuanLi_/status/2092616956432015754","https://x.com/OpenRouter/status/2092616758612172994","https://x.com/vllm_project/status/2092618480554332446","https://x.com/arena/status/2092622502757589440","https://x.com/SemiAnalysis_/status/2092623836671819852","https://x.com/ZixuanLi_/status/2092619063885218248"]
summary: "The model Z.ai had previously previewed as Ox Alpha is now available across its official platforms, giving the GLM-5 line its first natively multimodal release. Z.ai said GLM-5.3-Flash has a 1M token context window, 320B total parameters with 18B active, is released under the MIT license, and that the earlier Ox Alpha deployment ran entirely on Chinese AI chips. OpenRouter said Ox Alpha processed over 20 trillion tokens in six days before the reveal, making it the biggest model on its platform. Z.ai said the official release improves performance and stability, priced it at $0.15 per 1M input tokens and $0.50 per 1M output tokens, and is offering a 50% discount for the next two weeks. Arena later placed it around No. 5 on Code Arena WebDev in an early AutoEval ranking."
---

# Z\.ai Releases GLM\-5\.3\-Flash After Ox Alpha Processed 20T Tokens in 6 Days

> [Open the canonical story](<https://aidr.today/f1fbcda6?lang=en>)

**Published:** 2026-08-26T19:04:30.000Z
**Topics:** llm, open\-source, inference

## Summary

The model Z\.ai had previously previewed as Ox Alpha is now available across its official platforms, giving the GLM\-5 line its first natively multimodal release\. Z\.ai said GLM\-5\.3\-Flash has a 1M token context window, 320B total parameters with 18B active, is released under the MIT license, and that the earlier Ox Alpha deployment ran entirely on Chinese AI chips\. OpenRouter said Ox Alpha processed over 20 trillion tokens in six days before the reveal, making it the biggest model on its platform\. Z\.ai said the official release improves performance and stability, priced it at $0\.15 per 1M input tokens and $0\.50 per 1M output tokens, and is offering a 50% discount for the next two weeks\. Arena later placed it around No\. 5 on Code Arena WebDev in an early AutoEval ranking\.

## Sources

- [Story source](<https://huggingnews.com/ai/zai-releases-glm-53-flash-after-ox-alpha-processed-20t-tokens-in-6-days-262155ed>)
- [Story source](<https://x.com/Zai_org/status/2092616204787626030>)
- [Story source](<https://x.com/ZixuanLi_/status/2092616956432015754>)
- [Supporting source](<https://x.com/OpenRouter/status/2092616758612172994>)
- [Supporting source](<https://x.com/vllm_project/status/2092618480554332446>)
- [Supporting source](<https://x.com/arena/status/2092622502757589440>)
- [Supporting source](<https://x.com/SemiAnalysis_/status/2092623836671819852>)
- [Supporting source](<https://x.com/ZixuanLi_/status/2092619063885218248>)

