MiMo-V2.6 Pro, a model with 1.02T total parameters and 42B active parameters, is undergoing a Reinforcement Learning training run broadcast live via a public dashboard. The stream displays real-time internal metrics, including a daily operating cost of $493,000 and a $70,000 cost per step for the Pro version. A corresponding MiMo V2.6 Flash model, which contains 309B total parameters and 15B active parameters, is being trained at a cost of $247,000 per day.
The training process scales compute at approximately 2 billion tokens per step using 1,568 prompts and 16 rollouts. Xiaomi's team plans to open source the technical details over the next several weeks, with the final models hosted on Hugging Face. Other publicized live runs include the Marin 535B-A23B model, which has processed 5T of 18T tokens in one month, and a separate high-cost run by the developers of MOPD.
Key sources
- SOURCE@percyliang“Here's how Marin 535B-A23B is doing today”x.com
- SUPPORT@gauravisnotme“They are at about 5T/18T tokens in a month”x.com
- SUPPORT@madiator“Those behind MOPD are livestreaming a very expensive training run!”x.com
- SUPPORT@eliebakouch“literally livestreaming the RL training run of Mimo V2.6 Pro (1T, 42B active) and Flash (309B, 15B active)”x.com
- SUPPORT@eliebakouch“Mimo V2.6 Pro (1.02T total, 42B active param) 493k$ / day”x.com
- SUPPORT@omarsar0“$1M+ so far”x.com
- SUPPORT@natolambert“One of the coolest at-scale RL resources made public yet!”x.com
- SUPPORT@andrewcurran_“We believe RL is one of the most scalable and efficient paths toward self-improvement.”x.com