The computational demands of agentic inference are at the center of a new design released by Xiaomi for its next-generation AI models. The HySparse2 architecture, which will serve as the core of MiMo-V3, achieves 5.02x lower prefill FLOPs and a 4.5x smaller KV cache at 1M tokens compared to the Hybrid SWA architecture used in MiMo-V2.6. This update also improves long-context retrieval as measured by MRCRv2 and RULER-v2 scores while lowering AgentPPL and LongPPL. The design shifts to optimize for workloads where short actions return long observations that require frequent prefilling of a growing context. To reduce overhead, HySParse2 utilizes 'KV Bridging' to build cross-decoder K/V from self-decoder hidden states and 'KV Reuse' for sparse layers. Additional changes include moving to token-level selection and implementing a forced window of recent tokens to allow local and global tokens to share a single KV cache.

Sign in to suggest edits

Key sources

  1. SOURCEmarketbrief.now
  2. SOURCE@_luofuli“Agentic inference is a very different workload. Each round, a short action can return a long observation that needs to be prefilled, while the context keeps growing.”x.com
  3. SUPPORT@teortaxestex“HySparse went the way of NSA and MoBA, a published research idea”x.com
  4. SOURCEhuggingnewshuggingnews.com
Markdown