GGUF files are now available for Z.ai's GLM-5.3-Flash, the model previously previewed as Ox Alpha. Unsloth said the model can run in 3-bit form on systems with 128GB of RAM. Daniel Han Chen said a 4-bit version retains 93% accuracy and runs on a 256GB Mac or two DGX Sparks.

Z.ai introduced GLM-5.3-Flash this week under the MIT license with a 1M-token context window and said the model had been running entirely on Chinese AI chips. The model delivered nearly 20% weekly token share and ranked No. 1 on OpenRouter, which later said Ox Alpha processed over 20 trillion tokens in six days.

Sign in to suggest edits

Key sources

  1. SOURCE@unslothai“GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.”x.com
  2. SUPPORT@hesamation“Delivered nearly 20% weekly token share (no. 1) on OpenRouter.”x.com
  3. SUPPORT@zai_org“GLM-5.3-Flash can now be run locally! ✨”x.com
  4. SUPPORT@danielhanchen“It runs perfectly on a 256GB Mac or two DGX Sparks.”x.com
  5. SUPPORT@nvidiaai“881 tok/s C=64”x.com
  6. SOURCE@zai_org“GLM-5.3’s weights will be released tomorrow.”x.com
  7. SUPPORT@jietang“Natively multimodal with a 1M-token context…”x.com
  8. SUPPORT@ollama“it's available -- but rolling out, so performance isn't fully there yet!!”x.com
Markdown