---
format: "aidr-story-markdown/v1"
id: "9a8a1c36ebccf1015bc86c891e68cef657449a64a64fa5e41e00e0d0f064daca"
canonical_url: "https://aidr.today/9a8a1c36?lang=en"
title: "Berkeley and MIT Launch FreeToken to Run 753B AI Model on Single GPU 4x Faster Than Ollama"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-23T16:40:33.000Z"
category: "Infra"
topics: ["open-source","inference","moe","deepseek","qwen","chips"]
source_urls: ["https://huggingnews.com/ai/berkeley-and-mit-launch-freetoken-to-run-753b-ai-model-on-single-gpu-4x-d7e730fd","https://x.com/algo_diver/status/2091537643431682406","https://x.com/Huintellimance/status/2091541360918602033"]
summary: "An open source inference engine developed by academic researchers allows massive Mixture of Experts (MoE) models to run on consumer-grade hardware at interactive speeds. FreeToken enables a single RTX PRO 6000 GPU to serve the 753B parameter GLM-5.2 model at 14.9 tok/s and a 32GB desktop to run the 284B parameter DeepSeek-V4-Flash at 22 tok/s. An 8GB laptop GPU can run Qwen3.6 35B at 39.3 tok/s, with performance 2 to 4 times faster than Ollama on identical hardware using the same model weights. The software uses bandwidth adaptive execution to treat the GPU, CPU, RAM, and storage as a single elastic platform, dynamically mapping computation based on real time interconnect speeds. This approach optimizes MoE models like DeepSeek-V4-Flash, which only activate 13B of its 284B parameters per token. Developed by a team including Ion Stoica and Matei Zaharia, FreeToken is released under an Apache 2.0 license."
---

# Berkeley and MIT Launch FreeToken to Run 753B AI Model on Single GPU 4x Faster Than Ollama

> [Open the canonical story](<https://aidr.today/9a8a1c36?lang=en>)

**Published:** 2026-08-23T16:40:33.000Z
**Category:** Infra
**Topics:** open\-source, inference, moe, deepseek, qwen, chips

## Summary

An open source inference engine developed by academic researchers allows massive Mixture of Experts \(MoE\) models to run on consumer\-grade hardware at interactive speeds\. FreeToken enables a single RTX PRO 6000 GPU to serve the 753B parameter GLM\-5\.2 model at 14\.9 tok/s and a 32GB desktop to run the 284B parameter DeepSeek\-V4\-Flash at 22 tok/s\. An 8GB laptop GPU can run Qwen3\.6 35B at 39\.3 tok/s, with performance 2 to 4 times faster than Ollama on identical hardware using the same model weights\. The software uses bandwidth adaptive execution to treat the GPU, CPU, RAM, and storage as a single elastic platform, dynamically mapping computation based on real time interconnect speeds\. This approach optimizes MoE models like DeepSeek\-V4\-Flash, which only activate 13B of its 284B parameters per token\. Developed by a team including Ion Stoica and Matei Zaharia, FreeToken is released under an Apache 2\.0 license\.

## Sources

- [Story source](<https://huggingnews.com/ai/berkeley-and-mit-launch-freetoken-to-run-753b-ai-model-on-single-gpu-4x-d7e730fd>)
- [Story source](<https://x.com/algo_diver/status/2091537643431682406>)
- [Supporting source](<https://x.com/Huintellimance/status/2091541360918602033>)

