An open source inference engine developed by academic researchers allows massive Mixture of Experts (MoE) models to run on consumer-grade hardware at interactive speeds. FreeToken enables a single RTX PRO 6000 GPU to serve the 753B parameter GLM-5.2 model at 14.9 tok/s and a 32GB desktop to run the 284B parameter DeepSeek-V4-Flash at 22 tok/s. An 8GB laptop GPU can run Qwen3.6 35B at 39.3 tok/s, with performance 2 to 4 times faster than Ollama on identical hardware using the same model weights.

The software uses bandwidth adaptive execution to treat the GPU, CPU, RAM, and storage as a single elastic platform, dynamically mapping computation based on real time interconnect speeds. This approach optimizes MoE models like DeepSeek-V4-Flash, which only activate 13B of its 284B parameters per token. Developed by a team including Ion Stoica and Matei Zaharia, FreeToken is released under an Apache 2.0 license.

Sign in to suggest edits

Key sources

  1. SOURCE@algo_diver“UC Berkeley + MIT researchers open sourced a new inference engine”x.com
  2. SUPPORT@huintellimance“this is the biggest jump in edge inference since llama.cpp”x.com
Markdown