---
format: "aidr-story-markdown/v1"
id: "9f46901ad840a76cdbd83ffeaa19c8f37b6627c9da74c28f1047a687d9bcd92a"
canonical_url: "https://aidr.today/9f46901a?lang=en"
title: "Z.ai Reveals Ox Alpha as GLM-5.3-Flash, Biggest Model Yet on OpenRouter"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-28T03:00:29.000Z"
category: "Releases"
topics: ["llm","open-source","inference","chips"]
source_urls: ["https://huggingnews.com/ai/zai-reveals-ox-alpha-as-glm-53-flash-biggest-model-yet-on-openrouter-090aebe0","https://x.com/iScienceLuvr/status/2093043989750464889","https://x.com/TheTuringPost/status/2093098170423431604","https://x.com/prz_chojecki/status/2093040351632032124","https://x.com/theo/status/2093078228491731177","https://x.com/baseten/status/2093086722196172825"]
summary: "Developers probing a free mystery model across AI services were testing an early build of a new Z.ai system that the company later released with better performance and stability. Z.ai said the model, known publicly as Ox Alpha, was GLM-5.3-Flash, the first multimodal model in its GLM-5 series, released under the MIT license with a 1M token context window and 320B total parameters, 18B active at inference. OpenRouter said Ox Alpha became the biggest model ever on its marketplace, processing more than 20T tokens in 6 days. Z.ai put GLM-5.3-Flash on a two-week 50% discount through its API, lowering prices to $0.075 per 1M input tokens and $0.25 per 1M output tokens, and said the preview ran entirely on Chinese AI chips. In later testing, GLM-5.3-Flash ranked second on ErdosBench behind GPT-5.6 Sol xhigh."
---

# Z\.ai Reveals Ox Alpha as GLM\-5\.3\-Flash, Biggest Model Yet on OpenRouter

> [Open the canonical story](<https://aidr.today/9f46901a?lang=en>)

**Published:** 2026-08-28T03:00:29.000Z
**Category:** Releases
**Topics:** llm, open\-source, inference, chips

## Summary

Developers probing a free mystery model across AI services were testing an early build of a new Z\.ai system that the company later released with better performance and stability\. Z\.ai said the model, known publicly as Ox Alpha, was GLM\-5\.3\-Flash, the first multimodal model in its GLM\-5 series, released under the MIT license with a 1M token context window and 320B total parameters, 18B active at inference\. OpenRouter said Ox Alpha became the biggest model ever on its marketplace, processing more than 20T tokens in 6 days\. Z\.ai put GLM\-5\.3\-Flash on a two\-week 50% discount through its API, lowering prices to $0\.075 per 1M input tokens and $0\.25 per 1M output tokens, and said the preview ran entirely on Chinese AI chips\. In later testing, GLM\-5\.3\-Flash ranked second on ErdosBench behind GPT\-5\.6 Sol xhigh\.

## Sources

- [Story source](<https://huggingnews.com/ai/zai-reveals-ox-alpha-as-glm-53-flash-biggest-model-yet-on-openrouter-090aebe0>)
- [Story source](<https://x.com/iScienceLuvr/status/2093043989750464889>)
- [Supporting source](<https://x.com/TheTuringPost/status/2093098170423431604>)
- [Supporting source](<https://x.com/prz_chojecki/status/2093040351632032124>)
- [Supporting source](<https://x.com/theo/status/2093078228491731177>)
- [Supporting source](<https://x.com/baseten/status/2093086722196172825>)

