Build1 publisher3 min readPublished
Xiaomi's live training dashboard shows MiMo 2.6 Pro costing $20,500 an hour across seven restarts
Xiaomi's public dashboard put the MiMo 2.6 Pro reinforcement-learning run at $1.05 million after about 51 hours, roughly $20,500 an hour. Its restart notes and token count give other teams an all-in reference for pricing their own RL runs.
The Engineer · Build desk

What happened
- Xiaomi is streaming two unreleased runs: MiMo 2.6 Pro, about 1 trillion parameters with 42 billion active, and 2.6 Flash, 309 billion with 15 billion active.
- Each Pro step samples 1,568 prompts and generates 16 attempts at each, 25,088 sequences that run in sandboxes and get graded.
- At step 15 the sampler drew from 24 sources in five categories, and 1,061 of the 1,568 prompts, about 68 percent, were code.
- The dashboard's mean pass rate for Pro rose from 0.565 at step 1 to 0.615 at step 14.
- Xiaomi's own DeepSWE score rose from 58.4 to 65.8, and Hacker News commenters questioned the gain as possible contamination.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost If the counter is a flat wall-clock rate, failed hours bill like training hours, so an RL budget modelled on this run has to price downtime from launch onward.
- capability Teams sizing their own post-training spend now have a published per-token cost from a frontier-scale RL run to test their estimates against.
- constraint Xiaomi's DeepSWE gain cannot be counted as evidence of coding ability until someone shows the training sources do not contain DeepSWE's tasks.
At $5.71 a second, $1,048,236 is 183,579 seconds of spend, or 51.0 hours [1]. The Pro run started at 10:32 UTC on September 15 [5]. A dev.to write-up read the figures from the dashboard's JSON endpoints between 13:20 and 13:45 UTC on September 17 [4], which puts the reading 50.8 to 51.2 hours after launch [2]. So the counter matches a flat rate applied to wall-clock time since the start, with no pause for the seven restarts [3]. If that holds, an hour spent restarting costs the same as an hour spent training.
One restart every 7.3 hours or so is this run's failure rate [3], and the notices name the causes. At 21:14 UTC on September 16 the operators wrote that "the mimo-v2.6-pro run is restarting due to a vram issue on one node." [11] At 12:20 UTC the next day they wrote that "there was a network connectivity issue between the pro training cluster and the grader deployment." [13] The write-up's author wrote that "One bad GPU node can stall the whole synchronous job," and noted that the grader runs as a separate service [14]. The Flash run's failure was harder to see. It restarted from step 15 because "a type of infra error on one of datasets was not correctly detected over the past ~3 hours." [12]
The 12:20 notice also changed the training data. The operators "removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs." [13] The headline line averages pass and fail over whatever the sampler drew [9]. Steps after the removal sample from a pool with the cyber data taken out [7]. A pass-rate rise across that boundary mixes learning with a change in what is being graded.
The token count gives a second planning unit. By Thursday the Pro run had trained on 30.2 billion tokens using almost two million sandboxes [8]. Spread over those tokens, the counter comes to about $34.70 per million trained tokens, downtime included [4]. For that figure to carry over to another team, several things have to hold. The model needs an active parameter count near Pro's [1]. Rollouts need sandbox grading at a similar number of attempts per prompt, and similar length [6]. And the restart rate has to sit near Xiaomi's.
Xiaomi measured the DeepSWE gain itself, and it came to 7.4 points [5]. Most of every batch is code [7], so a coding benchmark is where real learning and leaked test tasks would both show up first. Ruling out contamination means matching the code sources against DeepSWE's tasks. The write-up does not report that Xiaomi has run that check.
I think the restart notes are the most useful thing on the page for an outside team. Posting them beside a running cost is good practice. Xiaomi did not announce any of it [15]. The community found the stream, and at 13:40 UTC on Thursday the latest post on @XiaomiMiMo's X account was still a September 8 notice about an invite-only desktop app beta [15].
What to watch
- Whether MiMo 2.6 ships open-weight under an MIT license as MiMo-V2.5-Pro did; open weights would let outsiders run their own DeepSWE overlap checks.
- Whether Xiaomi publishes the contents of the 24 data sources or a DeepSWE contamination check.
- Whether the Pro run's restart rate falls below roughly one every seven hours as training continues.