---
format: "aidr-story-markdown/v1"
id: "12196500c013ecfecdc0d43607d35d3cb30e23d30553134ea11a298b69cd5d47"
canonical_url: "https://aidr.today/12196500?lang=en"
title: "Unsloth Releases 3 Bit GLM-5.3-Flash for 128GB RAM After Reaching No. 1 on OpenRouter"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-27T17:02:04.000Z"
category: "Releases"
topics: ["open-source","llm","inference","fine-tuning"]
source_urls: ["https://huggingnews.com/ai/update-unsloth-releases-3-bit-glm-53-flash-for-128gb-ram-after-reaching-7cea2e2d","https://x.com/UnslothAI/status/2092986464196002094","https://x.com/Hesamation/status/2092960390615298444","https://x.com/Zai_org/status/2092997479318880513","https://x.com/danielhanchen/status/2092996385302094189","https://x.com/NVIDIAAI/status/2092966605579755903","https://x.com/Zai_org/status/2092814169263218860","https://x.com/jietang/status/2092840334040379779"]
summary: "GGUF files are now available for Z.ai's GLM-5.3-Flash, the model previously previewed as Ox Alpha. Unsloth said the model can run in 3-bit form on systems with 128GB of RAM. Daniel Han Chen said a 4-bit version retains 93% accuracy and runs on a 256GB Mac or two DGX Sparks. Z.ai introduced GLM-5.3-Flash this week under the MIT license with a 1M-token context window and said the model had been running entirely on Chinese AI chips. The model delivered nearly 20% weekly token share and ranked No. 1 on OpenRouter, which later said Ox Alpha processed over 20 trillion tokens in six days."
---

# Unsloth Releases 3 Bit GLM\-5\.3\-Flash for 128GB RAM After Reaching No\. 1 on OpenRouter

> [Open the canonical story](<https://aidr.today/12196500?lang=en>)

**Published:** 2026-08-27T17:02:04.000Z
**Category:** Releases
**Topics:** open\-source, llm, inference, fine\-tuning

## Summary

GGUF files are now available for Z\.ai's GLM\-5\.3\-Flash, the model previously previewed as Ox Alpha\. Unsloth said the model can run in 3\-bit form on systems with 128GB of RAM\. Daniel Han Chen said a 4\-bit version retains 93% accuracy and runs on a 256GB Mac or two DGX Sparks\. Z\.ai introduced GLM\-5\.3\-Flash this week under the MIT license with a 1M\-token context window and said the model had been running entirely on Chinese AI chips\. The model delivered nearly 20% weekly token share and ranked No\. 1 on OpenRouter, which later said Ox Alpha processed over 20 trillion tokens in six days\.

## Sources

- [Story source](<https://huggingnews.com/ai/update-unsloth-releases-3-bit-glm-53-flash-for-128gb-ram-after-reaching-7cea2e2d>)
- [Story source](<https://x.com/UnslothAI/status/2092986464196002094>)
- [Supporting source](<https://x.com/Hesamation/status/2092960390615298444>)
- [Supporting source](<https://x.com/Zai_org/status/2092997479318880513>)
- [Supporting source](<https://x.com/danielhanchen/status/2092996385302094189>)
- [Supporting source](<https://x.com/NVIDIAAI/status/2092966605579755903>)
- [Story source](<https://x.com/Zai_org/status/2092814169263218860>)
- [Supporting source](<https://x.com/jietang/status/2092840334040379779>)

