Skip to content

Build1 publisher3 min readPublished

Xiaomi's live training dashboard shows MiMo 2.6 Pro costing $20,500 an hour across seven restarts

Xiaomi's public dashboard put the MiMo 2.6 Pro reinforcement-learning run at $1.05 million after about 51 hours, roughly $20,500 an hour. Its restart notes and token count give other teams an all-in reference for pricing their own RL runs.

The Engineer · Build desk

Illustration accompanying Xiaomi's live training dashboard shows MiMo 2.6 Pro costing $20,500 an hour across seven restarts

What happened

  • Xiaomi is streaming two unreleased runs: MiMo 2.6 Pro, about 1 trillion parameters with 42 billion active, and 2.6 Flash, 309 billion with 15 billion active.
  • Each Pro step samples 1,568 prompts and generates 16 attempts at each, 25,088 sequences that run in sandboxes and get graded.
  • At step 15 the sampler drew from 24 sources in five categories, and 1,061 of the 1,568 prompts, about 68 percent, were code.
  • The dashboard's mean pass rate for Pro rose from 0.565 at step 1 to 0.615 at step 14.
  • Xiaomi's own DeepSWE score rose from 58.4 to 65.8, and Hacker News commenters questioned the gain as possible contamination.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost If the counter is a flat wall-clock rate, failed hours bill like training hours, so an RL budget modelled on this run has to price downtime from launch onward.
  • capability Teams sizing their own post-training spend now have a published per-token cost from a frontier-scale RL run to test their estimates against.
  • constraint Xiaomi's DeepSWE gain cannot be counted as evidence of coding ability until someone shows the training sources do not contain DeepSWE's tasks.

At $5.71 a second, $1,048,236 is 183,579 seconds of spend, or 51.0 hours [1]. The Pro run started at 10:32 UTC on September 15 [5]. A dev.to write-up read the figures from the dashboard's JSON endpoints between 13:20 and 13:45 UTC on September 17 [4], which puts the reading 50.8 to 51.2 hours after launch [2]. So the counter matches a flat rate applied to wall-clock time since the start, with no pause for the seven restarts [3]. If that holds, an hour spent restarting costs the same as an hour spent training.

One restart every 7.3 hours or so is this run's failure rate [3], and the notices name the causes. At 21:14 UTC on September 16 the operators wrote that "the mimo-v2.6-pro run is restarting due to a vram issue on one node." [11] At 12:20 UTC the next day they wrote that "there was a network connectivity issue between the pro training cluster and the grader deployment." [13] The write-up's author wrote that "One bad GPU node can stall the whole synchronous job," and noted that the grader runs as a separate service [14]. The Flash run's failure was harder to see. It restarted from step 15 because "a type of infra error on one of datasets was not correctly detected over the past ~3 hours." [12]

The 12:20 notice also changed the training data. The operators "removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs." [13] The headline line averages pass and fail over whatever the sampler drew [9]. Steps after the removal sample from a pool with the cyber data taken out [7]. A pass-rate rise across that boundary mixes learning with a change in what is being graded.

The token count gives a second planning unit. By Thursday the Pro run had trained on 30.2 billion tokens using almost two million sandboxes [8]. Spread over those tokens, the counter comes to about $34.70 per million trained tokens, downtime included [4]. For that figure to carry over to another team, several things have to hold. The model needs an active parameter count near Pro's [1]. Rollouts need sandbox grading at a similar number of attempts per prompt, and similar length [6]. And the restart rate has to sit near Xiaomi's.

Xiaomi measured the DeepSWE gain itself, and it came to 7.4 points [5]. Most of every batch is code [7], so a coding benchmark is where real learning and leaked test tasks would both show up first. Ruling out contamination means matching the code sources against DeepSWE's tasks. The write-up does not report that Xiaomi has run that check.

I think the restart notes are the most useful thing on the page for an outside team. Posting them beside a running cost is good practice. Xiaomi did not announce any of it [15]. The community found the stream, and at 13:40 UTC on Thursday the latest post on @XiaomiMiMo's X account was still a September 8 notice about an invite-only desktop app beta [15].

What to watch

  • Whether MiMo 2.6 ships open-weight under an MIT license as MiMo-V2.5-Pro did; open weights would let outsiders run their own DeepSWE overlap checks.
  • Whether Xiaomi publishes the contents of the 24 data sources or a DeepSWE contamination check.
  • Whether the Pro run's restart rate falls below roughly one every seven hours as training continues.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories