---
format: "aidr-story-markdown/v1"
id: "f1dca09efa5746ae7cfd7149212ef989f7919d43df80469d0d831fa076809d3e"
canonical_url: "https://aidr.today/f1dca09e?lang=en"
title: "Xiaomi Cuts MiMo-V3 Prefill Compute 5.02x With HySParse2 Architecture"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-23T17:17:53.000Z"
category: "Research"
topics: ["xiaomi","inference","agent","chips","infra"]
source_urls: ["https://marketbrief.now/ai/update-xiaomi-cuts-mimo-v3-prefill-compute-502x-with-hysparse2-architect-1fc708bb","https://x.com/_LuoFuli/status/2102766365190901957","https://x.com/teortaxesTex/status/2102775394206179335","https://huggingnews.com/ai/update-xiaomi-cuts-mimo-v3-prefill-compute-502x-with-hysparse2-architect-1fc708bb"]
summary: "The computational demands of agentic inference are at the center of a new design released by Xiaomi for its next-generation AI models. The HySparse2 architecture, which will serve as the core of MiMo-V3, achieves 5.02x lower prefill FLOPs and a 4.5x smaller KV cache at 1M tokens compared to the Hybrid SWA architecture used in MiMo-V2.6. This update also improves long-context retrieval as measured by MRCRv2 and RULER-v2 scores while lowering AgentPPL and LongPPL. The design shifts to optimize for workloads where short actions return long observations that require frequent prefilling of a growing context. To reduce overhead, HySParse2 utilizes 'KV Bridging' to build cross-decoder K/V from self-decoder hidden states and 'KV Reuse' for sparse layers. Additional changes include moving to token-level selection and implementing a forced window of recent tokens to allow local and global tokens to share a single KV cache."
---

# Xiaomi Cuts MiMo\-V3 Prefill Compute 5\.02x With HySParse2 Architecture

> [Open the canonical story](<https://aidr.today/f1dca09e?lang=en>)

**Published:** 2026-09-23T17:17:53.000Z
**Category:** Research
**Topics:** xiaomi, inference, agent, chips, infra

## Summary

The computational demands of agentic inference are at the center of a new design released by Xiaomi for its next\-generation AI models\. The HySparse2 architecture, which will serve as the core of MiMo\-V3, achieves 5\.02x lower prefill FLOPs and a 4\.5x smaller KV cache at 1M tokens compared to the Hybrid SWA architecture used in MiMo\-V2\.6\. This update also improves long\-context retrieval as measured by MRCRv2 and RULER\-v2 scores while lowering AgentPPL and LongPPL\. The design shifts to optimize for workloads where short actions return long observations that require frequent prefilling of a growing context\. To reduce overhead, HySParse2 utilizes 'KV Bridging' to build cross\-decoder K/V from self\-decoder hidden states and 'KV Reuse' for sparse layers\. Additional changes include moving to token\-level selection and implementing a forced window of recent tokens to allow local and global tokens to share a single KV cache\.

## Sources

- [Story source](<https://marketbrief.now/ai/update-xiaomi-cuts-mimo-v3-prefill-compute-502x-with-hysparse2-architect-1fc708bb>)
- [Story source](<https://x.com/_LuoFuli/status/2102766365190901957>)
- [Supporting source](<https://x.com/teortaxesTex/status/2102775394206179335>)
- [Story source](<https://huggingnews.com/ai/update-xiaomi-cuts-mimo-v3-prefill-compute-502x-with-hysparse2-architect-1fc708bb>)

