Xiaomi unveiled the AI Cube, a personal inference hardware prototype designed to run large language models locally. The system integrates three proprietary chips—Xring O3, O100, and D100—to support the simultaneous deployment of 120B and 3B parameter models. The O100 accelerator uses 6nm wafer-on-wafer NPU-DRAM bonding to provide 1.22TB/s of near-memory bandwidth, while the D100 processor, originally designed for smart driving, supports 160GB of memory and local deployment of models up to 200B parameters.

The prototype maintains a 150W continuous power release and incorporates the Xring O3 high-end SoC, which is manufactured by TSMC using a 3nm process. While the O3 chip is targeted for Xiaomi's next flagship foldable phones, its inclusion in the AI Cube moves the company into the desktop AI workstation market to compete with products like Nvidia's DGX Spark and Apple's Mac Studio.

Sign in to suggest edits

Key sources

  1. SOURCE@jun_song“runs a 120B and a 3B model at the same time”x.com
  2. SUPPORT@maxforai“D100 is Xiaomi's smart driving high-performance AI chip, supporting up to 160GB of memory and 200B large model local deployment”x.com
  3. SUPPORT@teortaxestex“6nm, Wafer-on-wafer NPU-DRAM bonding with 1.4µm pitch”x.com
  4. SUPPORT@tonyjzhou“Decoding re-reads active weights every token, so local inference is bandwidth-bound, not FLOPs-bound”x.com
  5. SUPPORT@hesamation“Xiaomi's new prototype looks like a direct competitor to DGX Spark”x.com
  6. SOURCE@reuters“partners with TSMC for production”x.com
  7. SUPPORT@ftr_investors“200K–300K units targeted — reducing its reliance on Qualcomm and MediaTek”x.com
  8. SUPPORT@firstsquawk“Xring D100 3nm smart driving processor”x.com
Markdown