← Back to live feed1 story
Xiaomi's Xring O100 chip uses 3D wafer-on-wafer stacking to enable on-device large language model inference for smartphones, cars, and robots. The 6nm accelerator delivers 1.22 TB/s of near-memory bandwidth, which is 16 times that of a flagship phone's LPDDR5X memory. In lab tests, the hardware processed a 3B parameter model at 330 tokens per second.
The chip employs Hybrid Bonding to shrink the bonding pitch from 50 μm to 1.4 μm, creating 2.58 million physical pathways via through-silicon vias. Xiaomi is deploying the O100 in an AI Cube prototype alongside Xring O3 and D100 processors to run models up to 120B parameters at 150W sustained power. The chip has completed silicon validation and is scheduled for commercial use in 2027.
Sign in to suggest edits
Key sources
- SOURCE@neil_shah“Bonding pitch shrinks from ~50 μm all the way to 1.4 μm”x.com
- SUPPORT@jiemian_news“can run Xiaomi’s MiMo 3B model at 330 tokens/s”x.com
- SUPPORT@zephyr_z9“This spec is extremely similar to Huawei's Logic Folding”x.com
- SOURCE@jun_song“runs a 120B and a 3B model at the same time”x.com
- SUPPORT@maxforai“D100 is Xiaomi's smart driving high-performance AI chip, supporting up to 160GB of memory and 200B large model local deployment”x.com
- SUPPORT@teortaxestex“6nm, Wafer-on-wafer NPU-DRAM bonding with 1.4µm pitch”x.com
- SUPPORT@tonyjzhou“Decoding re-reads active weights every token, so local inference is bandwidth-bound, not FLOPs-bound”x.com
- SUPPORT@hesamation“Xiaomi's new prototype looks like a direct competitor to DGX Spark”x.com