Xiaomi MiMo-V2.6 Training Costs Top USD 1.28 Million, Luo Fuli Says
Luo Fuli, who leads Xiaomi's MiMo large-model team, disclosed that MiMo-V2.6 is midway through reinforcement learning training, with cumulative costs exceeding USD 1.28 million. MiMo-V2.6-Pro has run for one day and 19 hours at USD 890,000, averaging over USD 20,000 per hour, while MiMo-V2.6-Flash has run for one day and 14 hours at USD 397,000, averaging over USD 10,000 per hour. Details will be open-sourced in the coming weeks.
In the early hours of September 17, Luo Fuli, who leads Xiaomi's MiMo large-model team, said in a post on X that the team is currently studying how far reinforcement learning (RL) can go. Xiaomi's latest large model, MiMo-V2.6, is midway through its reinforcement learning run, and the team is scaling along three dimensions during this period: on compute, roughly 2 billion tokens per step, 1,568 prompts multiplied by 16 rollouts, and fully asynchronous; on environments and test frameworks, multi-task agentic RL that mixes multiple test frameworks in a single run; and on evaluation compute, agentic intra-group credit assignment with test cases and rubric-based rewards. Luo Fuli said the relevant details will be open-sourced progressively over the coming weeks.
Luo Fuli also shared a real-time training page for MiMo-V2.6, the first time outsiders have seen part of the training status of this generation of models during the reinforcement learning stage. The page displays training progress, token consumption, training costs, sample counts, and changes in several internal metrics, covering two versions, MiMo-V2.6-Pro and MiMo-V2.6-Flash.
The cumulative training cost of the two versions has now exceeded USD 1.28 million. MiMo-V2.6-Pro has trained for one day and 19 hours at a cost of USD 890,000, averaging more than USD 20,000 per hour; MiMo-V2.6-Flash has trained for one day and 14 hours at a cost of USD 397,000, averaging more than USD 10,000 per hour.
In March 2026, Lei Jun said in a post on Weibo that the trillion-parameter large model MiMo-V2-Pro ranked eighth globally on the Artificial Analysis comprehensive intelligence leaderboard for large models; by large-model brand ranking, it placed fifth globally, ahead of xAI Grok.
Luo Fuli completed her undergraduate studies in computer science at Beijing Normal University and pursued computational linguistics at Peking University for her master's degree. After graduating, she joined Alibaba DAMO Academy as a researcher in the Machine Intelligence Laboratory, where she was responsible for developing the multilingual pre-trained model VECO and drove the open-sourcing of the AliceMind project. In 2022, Luo Fuli joined High-Flyer Quant, the parent company of DeepSeek, to work on deep learning, and later served as a deep learning researcher at DeepSeek, participating in the development of models including DeepSeek-V2.
On November 12, 2025, Luo Fuli said in a post on WeChat Moments: "Intelligence will eventually move from language to the physical world. I am at Xiaomi MiMo, working with a group of researchers to build such a future, racing toward the AGI we envision." This was also the first time Luo Fuli formally announced that she had joined Xiaomi.
Why this event matters
The event has a measured impact on 2 industrys. The strongest current signal is positive for Artificial Intelligence, with intensity 60/100 and 70% confidence over a short term horizon.
Artificial Intelligence
- Direction
- positive
- Intensity
- 60
- Confidence
- 70%
- Horizon
- Short term
Cloud Services & Data Centres
- Direction
- positive
- Intensity
- 50
- Confidence
- 60%
- Horizon
- Medium term
Impact figures are analytical estimates that combine direction, intensity, confidence and event importance. They are not investment advice.