BuildWidely confirmed7 publishers3 min readPublished Updated
Xiaomi priced six days of reinforcement learning at $3.47 million
Xiaomi's $850,000 and $2.62 million cover 30 reinforcement-learning steps apiece on models it had already pretrained. Xiaomi published the environments and the RL code but not the 7,000-plus task datasets behind them.
The Engineer · Build desk

What happened
- Xiaomi released and open-sourced MiMo-V2.6-Pro and MiMo-V2.6-Flash, two natively omnimodal models, with Pro at 1.02 trillion total parameters and 42 billion activated per token.
- The release includes the technical report, the training environments and the RL code, short of the complete set of more than 7,000 training tasks.
- Hosted prices hold at V2.5 levels: Pro at $0.435 per million input tokens and $0.87 output, Flash at $0.14 and $0.28, with the UltraSpeed tier costing ten times Pro.
Why it matters
- constraint The published cost covers the RL stage on top of a base model Xiaomi already had. A team using $3.47 million as its planning figure is pricing only the last stage of the build.
- cost Running Pro yourself means distributed inference across multiple nodes and GPUs. Self-hosting lands on the capital budget, and the hosted token price stays on the operating one.
- decision The DeepSWE gains are Xiaomi's own unreproduced measurements. Anyone evaluating MiMo has to build a verifiable task suite and graders of their own before those numbers inform a procurement choice.
Both cost figures Xiaomi published attach to one stage of the work. Flash's run cost about $850,000 and Pro's about $2.62 million, and each covers 30 reinforcement-learning steps across roughly 750,000 trajectories finished in under six days [28][5]. Together that is $3.47 million [18]. Both figures are for the RL runs, with pretraining outside either account [27].
The batch shape shows where the money went. Each update used 1,568 prompts with 16 rollouts per prompt, and each step produced between 3.5 billion and 3.7 billion training tokens [29]. Thirty steps at 1,568 by 16 is 752,640 trajectories, which matches the roughly 750,000 reported [30]. Pro's run therefore paid about $3.48 per trajectory [19] and generated somewhere between 105 billion and 111 billion tokens end to end [31].
A binary test can tell a training system that an agent finished a task. It cannot separate a concise, reliable solution from a wasteful one that happened to pass. Grader design is the substantive part of what Xiaomi disclosed: its graders rank the successful attempts against one another and redistribute reward toward the higher-quality paths [15]. Coding, general agent, visual and cybersecurity tasks went into the same run instead of one model per category [14]. Xiaomi says it froze the router to limit drift and used adversarial evaluation, anomaly detection and verifier cross-checks against reward hacking [16].
Reproducing the run takes more than the released package. Xiaomi is open-sourcing the environment code and training recipes, minus the complete 7,000-plus task datasets [6]. Anyone repeating this needs their own verifiable tasks and their own pretrained base, so the $2.62 million is the cost of the last stage.
Serving is the other adoption cost. The Pro model card describes a sparse mixture-of-experts with 1.02 trillion total parameters and 42 billion activated per token [2]. Xiaomi's deployment recipe for Pro recommends distributed inference across multiple nodes and GPUs [17]. Most teams will compare hosted prices. Pro is $0.435 per million input tokens and $0.87 per million output, Flash $0.14 and $0.28, and UltraSpeed ten times Pro [7], which is $4.35 and $8.70 [22].
Pro scored 46.32 on version 4.3 of the Artificial Analysis Intelligence Index, which Xiaomi says puts it ahead of Kimi K3 and Qwen3.8 Max among the open models in that comparison [8]. The rest of the table is mixed. Pro trails several closed models on ProgramBench, Terminal Bench 4.0 and exploit-generation tests, and competes better on automation, visual coding and some long-horizon agent evaluations [24]. On DeepSWE v1.1, Xiaomi reports Pro moving from 58.4 to 72.57 and Flash from 48.8 to 65.68, gains of 14.17 and 16.88 points [10][23]. Those are Xiaomi's own evaluation, not independently reproduced [10]. For the deltas to transfer, your tasks have to be verifiable the way DeepSWE's are, and your graders have to agree with Xiaomi's about what a good trajectory looks like.
The MiMo team is led by Luo Fuli, the former DeepSeek researcher who joined Xiaomi to head it, VnExpress reported in November 2025 [20]. She published the final RL training runs live, internal metrics visible [21]. Xiaomi has not traditionally been counted among the six Chinese AI Tigers [26]. On what this means for closed APIs, @Yuchenj_UW claims frontier coding capability has plateaued since Opus 4.8 while open models keep closing the gap at 10 to 50 times lower cost [32]. @teortaxesTex argues the frontier has split into new higher tiers where internal and top closed models are still well ahead [33].
What to watch
- An independent reproduction of the DeepSWE v1.1 numbers, which Xiaomi ran on its own evaluation.
- Release of the complete 7,000-plus task datasets. Until they arrive, the published environments and recipes need someone else's tasks.
- Whether the 46.32 Intelligence Index placing survives the next index version or the next open-weights release.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence60
- Adoption25
- Hype gap+30
- Incentives65
- Confidence62
Perspective Coverage
7 publishers- Builder
- Builder 56%
- Operator
- Operator 23%
- Investor
- Investor 21%
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Xiaomi has released and open-sourced the MiMo-V2.6 series, two natively omnimodal models, MiMo-V2.6-Pro and MiMo-V2.6-Flash, for coding, visual tasks and computer use.
- [2]
The MiMo-V2.6-Pro model card describes a sparse mixture-of-experts architecture with 1.02 trillion total parameters and 42 billion activated for each token.
- [3]
The MiMo-V2.6-Flash model card lists 309 billion total parameters and 15 billion activated.
- [4]
Both MiMo-V2.6 models accept text, images, video and audio, and both have a claimed 1-million-token context window.
- [5]
The 30-step RL runs cost about $850,000 for Flash and about $2.62 million for Pro.
- [6]
Xiaomi will open-source the tooling including the environment code and training recipes, but the complete 7k+ task datasets have not yet been released.
- [7]
API prices stay at V2.5 levels, with Flash at $0.14 per million uncached input tokens and $0.28 for output, Pro at $0.435 and $0.87, and UltraSpeed costing ten times more.
- [8]
MiMo-V2.6-Pro scored 46.32 on the Artificial Analysis Intelligence Index v4.3, which Xiaomi says puts it ahead of Kimi K3 and Qwen3.8 Max as the highest-scoring open-source model in the comparison.
- [9]
Artificial Analysis says MiMo-V2.6-Pro debuts as the top open-weights model on its Intelligence Index (46), with 1.02T total / 42B active parameters and strong cost efficiency at $0.435/M input and $0.87/M output tokens.
- [10]
On DeepSWE v1.1, a held-out software-engineering test, Xiaomi reports Flash rising from 48.8 to 65.68 and Pro moving from 58.4 to 72.57; those results come from Xiaomi's own evaluation and have yet to be independently reproduced.
- [11]
The MiMo-V2.6 models are under the MIT license, according to @victormustar.
- [12]
Xiaomi is rolling out MiMo-V2.6-Pro-UltraSpeed, delivering up to 20x faster output speed at the same quality.
- [13]
Pro and Flash are available in AI Studio, MiMo Code, MiMo Desktop, Xiaomi's MiMo API Platform and OpenRouter.
- [14]
Xiaomi mixed coding, general agent, visual and cybersecurity tasks into the same run instead of training a separate model for each category.
- [15]
Xiaomi's MiMo team built graders that compare successful trajectories against one another: a binary test can tell a training system whether an agent completed a task, though it cannot distinguish a concise, reliable solution from a wasteful one that happened to pass, so Xiaomi's approach ranks successful attempts and redistributes the training reward toward higher-quality paths.
- [16]
Xiaomi froze the router to limit drift and used adversarial evaluation, anomaly detection and verifier cross-checks against reward hacking.
- [17]
Model weights let developers operate the checkpoints on their own infrastructure, although models of this size remain expensive to serve, and Xiaomi's deployment recipe for Pro recommends distributed inference across multiple nodes and GPUs.
- [18]
The two reported RL runs together cost about $3.47 million.
- [19]
Pro's reported run cost about $3.48 per trajectory.
- [20]
The technical work is led by Luo Fuli, the former DeepSeek researcher who joined Xiaomi to head the MiMo team, VnExpress reported in November 2025; she previously worked at Alibaba's DAMO Academy.
- [21]
Fuli Luo, a former DeepSeek engineer now at Xiaomi, started publishing the final RL training runs live, showing an abnormal amount of transparency in internal metrics.
- [22]
UltraSpeed pricing works out to $4.35 per million input tokens and $8.70 per million output tokens.
- [23]
The reported DeepSWE v1.1 gains are 14.17 points for Pro and 16.88 points for Flash.
- [24]
MiMo-V2.6-Pro trails several closed models on ProgramBench, Terminal Bench 4.0 and exploit-generation tests, and performs more competitively on automation, visual coding and some long-horizon agent evaluations.
- [25]
Xiaomi says Flash's average pass rate on training tasks increased 25% in relative terms, while Pro improved 12%.
- [26]
Xiaomi is not traditionally considered one of the six Chinese AI Tigers.
- [27]
The published cost figures are attributed to the 30-step reinforcement-learning production runs, and no pretraining cost is reported for either model.
- [28]
MiMo-V2.6-Pro and Flash each completed 30 reinforcement-learning steps across roughly 750,000 trajectories in fewer than six days, according to Xiaomi's account of the production run.
- [29]
Each update used 1,568 prompts with 16 rollouts per prompt, producing between 3.5 billion and 3.7 billion training tokens per step.
- [30]
The stated batch shape gives 752,640 trajectories over the run, consistent with the roughly 750,000 reported.
- [31]
The 30-step run generated roughly 105 billion to 111 billion training tokens in total.
- [32]
@Yuchenj_UW claims frontier coding capability has plateaued since Opus 4.8, while open-source models keep closing the gap at 10-50x lower cost.
- [33]
@teortaxesTex argues frontier has actually split into new higher tiers, with internal models and top closed models still well ahead.
Sources
7 independent publishers whose own reporting we read for this story.
- blog.vercel.comMiMo V2.6 models now available on AI Gateway
1 article · September 20, 2026
- huggingface.coXiaomiMiMo/MiMo-V2.6-Pro-RL · Hugging Face
1 article · September 22, 2026
- latent.space[AINews] Xiaomi MiMo-V2.6-Pro 1T-A42B: the new top Open Weights model, trained for $3M
1 article · September 21, 2026
- runtimewire.comXiaomi open-sources MiMo-V2.6 and the RL machinery behind it
1 article · September 21, 2026
- testingcatalog.comXiaomi open-sources MiMo-V2.6 Pro and Flash models
2 articles · September 21, 2026
- the-decoder.comXiaomi's affordable flagship AI leads the open models, and Anthropic says Claude helped get it there
2 articles · September 22, 2026
- thenewstack.io“Impressive level of openness”: Xiaomi goes way beyond the usual open-weight playbook with MiMo-V2.6
1 article · September 23, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.