The model Z.ai had previously previewed as Ox Alpha is now available across its official platforms, giving the GLM-5 line its first natively multimodal release. Z.ai said GLM-5.3-Flash has a 1M token context window, 320B total parameters with 18B active, is released under the MIT license, and that the earlier Ox Alpha deployment ran entirely on Chinese AI chips.

OpenRouter said Ox Alpha processed over 20 trillion tokens in six days before the reveal, making it the biggest model on its platform. Z.ai said the official release improves performance and stability, priced it at $0.15 per 1M input tokens and $0.50 per 1M output tokens, and is offering a 50% discount for the next two weeks. Arena later placed it around No. 5 on Code Arena WebDev in an early AutoEval ranking.

Sign in to suggest edits

Key sources

  1. SOURCE@zai_org“Previously previewed as Ox Alpha, running entirely on Chinese AI chips”x.com
  2. SOURCE@zixuanli_“The official release delivers stronger performance and significantly better stability.”x.com
  3. SUPPORT@openrouter“the biggest model ever on OpenRouter”x.com
  4. SUPPORT@vllm_project“320B total, 18B active, 45 layers where GLM-4.5 had 92.”x.com
  5. SUPPORT@arena“landed around #5 in the Code Arena: WebDev (#2 among open models) scoring 1634 (AutoEval).”x.com
  6. SUPPORT@semianalysis_“attaining hardware efficiency and per-token cost comparable to Nvidia GPUs.”x.com
  7. SUPPORT@zixuanli_“GLM-5.3-Flash is now 50% off through the official https://t.co/GTCIA6vbgh API for the next two weeks.”x.com
  8. SOURCE@zai_org“performs on par with Claude Opus 4.8”x.com
Markdown