← Back to live feed1 story
1
Unsloth Releases 3 Bit GLM-5.3-Flash for 128GB RAM After Reaching No. 1 on OpenRouter
Releases36d agoGGUF files are now available for Z.ai's GLM-5.3-Flash, the model previously previewed as Ox Alpha. Unsloth said the model can run in 3-bit form on systems with 128GB of RAM. Daniel Han Chen said a 4-bit version retains 93% accuracy and runs on a 256GB Mac or two DGX Sparks.
Z.ai introduced GLM-5.3-Flash this week under the MIT license with a 1M-token context window and said the model had been running entirely on Chinese AI chips. The model delivered nearly 20% weekly token share and ranked No. 1 on OpenRouter, which later said Ox Alpha processed over 20 trillion tokens in six days.
Sign in to suggest edits
Key sources
- SOURCE@unslothai“GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.”x.com
- SUPPORT@hesamation“Delivered nearly 20% weekly token share (no. 1) on OpenRouter.”x.com
- SUPPORT@zai_org“GLM-5.3-Flash can now be run locally! ✨”x.com
- SUPPORT@danielhanchen“It runs perfectly on a 256GB Mac or two DGX Sparks.”x.com
- SUPPORT@nvidiaai“881 tok/s C=64”x.com
- SOURCE@zai_org“GLM-5.3’s weights will be released tomorrow.”x.com
- SUPPORT@jietang“Natively multimodal with a 1M-token context…”x.com
- SUPPORT@ollama“it's available -- but rolling out, so performance isn't fully there yet!!”x.com