---
format: "aidr-story-markdown/v1"
id: "7cd40c0f9b9f4f579b61ffe8df7f2956154fdf776cc447f386d401464d62d8c4"
canonical_url: "https://aidr.today/7cd40c0f?lang=en"
title: "Xiaomi AI Cube Prototype Runs 120B Model Locally With 1.22TB/s Bandwidth"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-24T11:44:45.000Z"
category: "Infra"
topics: ["xiaomi","inference","chips","hardware"]
source_urls: ["https://huggingnews.com/ai/xiaomi-ai-cube-prototype-runs-120b-model-locally-with-122tbs-bandwidth-044af1c4","https://x.com/jun_song/status/2091795576384106901","https://x.com/MaxForAI/status/2091805969718415551","https://x.com/teortaxesTex/status/2091794713674186898","https://x.com/TonyJZhou/status/2091847724245151857","https://x.com/Hesamation/status/2091834625714712907","https://x.com/Reuters/status/2091780075465158983","https://x.com/ftr_investors/status/2091784143570883012"]
summary: "Xiaomi unveiled the AI Cube, a personal inference hardware prototype designed to run large language models locally. The system integrates three proprietary chips—Xring O3, O100, and D100—to support the simultaneous deployment of 120B and 3B parameter models. The O100 accelerator uses 6nm wafer-on-wafer NPU-DRAM bonding to provide 1.22TB/s of near-memory bandwidth, while the D100 processor, originally designed for smart driving, supports 160GB of memory and local deployment of models up to 200B parameters. The prototype maintains a 150W continuous power release and incorporates the Xring O3 high-end SoC, which is manufactured by TSMC using a 3nm process. While the O3 chip is targeted for Xiaomi's next flagship foldable phones, its inclusion in the AI Cube moves the company into the desktop AI workstation market to compete with products like Nvidia's DGX Spark and Apple's Mac Studio."
---

# Xiaomi AI Cube Prototype Runs 120B Model Locally With 1\.22TB/s Bandwidth

> [Open the canonical story](<https://aidr.today/7cd40c0f?lang=en>)

**Published:** 2026-08-24T11:44:45.000Z
**Category:** Infra
**Topics:** xiaomi, inference, chips, hardware

## Summary

Xiaomi unveiled the AI Cube, a personal inference hardware prototype designed to run large language models locally\. The system integrates three proprietary chips—Xring O3, O100, and D100—to support the simultaneous deployment of 120B and 3B parameter models\. The O100 accelerator uses 6nm wafer\-on\-wafer NPU\-DRAM bonding to provide 1\.22TB/s of near\-memory bandwidth, while the D100 processor, originally designed for smart driving, supports 160GB of memory and local deployment of models up to 200B parameters\. The prototype maintains a 150W continuous power release and incorporates the Xring O3 high\-end SoC, which is manufactured by TSMC using a 3nm process\. While the O3 chip is targeted for Xiaomi's next flagship foldable phones, its inclusion in the AI Cube moves the company into the desktop AI workstation market to compete with products like Nvidia's DGX Spark and Apple's Mac Studio\.

## Sources

- [Story source](<https://huggingnews.com/ai/xiaomi-ai-cube-prototype-runs-120b-model-locally-with-122tbs-bandwidth-044af1c4>)
- [Story source](<https://x.com/jun_song/status/2091795576384106901>)
- [Supporting source](<https://x.com/MaxForAI/status/2091805969718415551>)
- [Supporting source](<https://x.com/teortaxesTex/status/2091794713674186898>)
- [Supporting source](<https://x.com/TonyJZhou/status/2091847724245151857>)
- [Supporting source](<https://x.com/Hesamation/status/2091834625714712907>)
- [Story source](<https://x.com/Reuters/status/2091780075465158983>)
- [Supporting source](<https://x.com/ftr_investors/status/2091784143570883012>)

