Berkeley and MIT Launch FreeToken to Run 753B AI Model on Single GPU 4x Faster Than Ollama
Infra39d agoAn open source inference engine developed by academic researchers allows massive Mixture of Experts (MoE) models to run on consumer-grade hardware at interactive speeds. FreeToken enables a single RTX PRO 6000 GPU to serve the 753B parameter GLM-5.2 model at 14.9 tok/s and a 32GB desktop to run the 284B parameter DeepSeek-V4-Flash at 22 tok/s. An 8GB laptop GPU can run Qwen3.6 35B at 39.3 tok/s, with performance 2 to 4 times faster than Ollama on identical hardware using the same model weights.
The software uses bandwidth adaptive execution to treat the GPU, CPU, RAM, and storage as a single elastic platform, dynamically mapping computation based on real time interconnect speeds. This approach optimizes MoE models like DeepSeek-V4-Flash, which only activate 13B of its 284B parameters per token. Developed by a team including Ion Stoica and Matei Zaharia, FreeToken is released under an Apache 2.0 license.