Skip to content

BuildWidely confirmed7 publishers3 min readPublished Updated

Xiaomi priced six days of reinforcement learning at $3.47 million

Xiaomi's $850,000 and $2.62 million cover 30 reinforcement-learning steps apiece on models it had already pretrained. Xiaomi published the environments and the RL code but not the 7,000-plus task datasets behind them.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Xiaomi priced six days of reinforcement learning at $3.47 million
Generated illustration

What happened

  • Xiaomi released and open-sourced MiMo-V2.6-Pro and MiMo-V2.6-Flash, two natively omnimodal models, with Pro at 1.02 trillion total parameters and 42 billion activated per token.
  • The release includes the technical report, the training environments and the RL code, short of the complete set of more than 7,000 training tasks.
  • Hosted prices hold at V2.5 levels: Pro at $0.435 per million input tokens and $0.87 output, Flash at $0.14 and $0.28, with the UltraSpeed tier costing ten times Pro.

Why it matters

  • constraint The published cost covers the RL stage on top of a base model Xiaomi already had. A team using $3.47 million as its planning figure is pricing only the last stage of the build.
  • cost Running Pro yourself means distributed inference across multiple nodes and GPUs. Self-hosting lands on the capital budget, and the hosted token price stays on the operating one.
  • decision The DeepSWE gains are Xiaomi's own unreproduced measurements. Anyone evaluating MiMo has to build a verifiable task suite and graders of their own before those numbers inform a procurement choice.

Both cost figures Xiaomi published attach to one stage of the work. Flash's run cost about $850,000 and Pro's about $2.62 million, and each covers 30 reinforcement-learning steps across roughly 750,000 trajectories finished in under six days [28][5]. Together that is $3.47 million [18]. Both figures are for the RL runs, with pretraining outside either account [27].

The batch shape shows where the money went. Each update used 1,568 prompts with 16 rollouts per prompt, and each step produced between 3.5 billion and 3.7 billion training tokens [29]. Thirty steps at 1,568 by 16 is 752,640 trajectories, which matches the roughly 750,000 reported [30]. Pro's run therefore paid about $3.48 per trajectory [19] and generated somewhere between 105 billion and 111 billion tokens end to end [31].

A binary test can tell a training system that an agent finished a task. It cannot separate a concise, reliable solution from a wasteful one that happened to pass. Grader design is the substantive part of what Xiaomi disclosed: its graders rank the successful attempts against one another and redistribute reward toward the higher-quality paths [15]. Coding, general agent, visual and cybersecurity tasks went into the same run instead of one model per category [14]. Xiaomi says it froze the router to limit drift and used adversarial evaluation, anomaly detection and verifier cross-checks against reward hacking [16].

Reproducing the run takes more than the released package. Xiaomi is open-sourcing the environment code and training recipes, minus the complete 7,000-plus task datasets [6]. Anyone repeating this needs their own verifiable tasks and their own pretrained base, so the $2.62 million is the cost of the last stage.

Serving is the other adoption cost. The Pro model card describes a sparse mixture-of-experts with 1.02 trillion total parameters and 42 billion activated per token [2]. Xiaomi's deployment recipe for Pro recommends distributed inference across multiple nodes and GPUs [17]. Most teams will compare hosted prices. Pro is $0.435 per million input tokens and $0.87 per million output, Flash $0.14 and $0.28, and UltraSpeed ten times Pro [7], which is $4.35 and $8.70 [22].

Pro scored 46.32 on version 4.3 of the Artificial Analysis Intelligence Index, which Xiaomi says puts it ahead of Kimi K3 and Qwen3.8 Max among the open models in that comparison [8]. The rest of the table is mixed. Pro trails several closed models on ProgramBench, Terminal Bench 4.0 and exploit-generation tests, and competes better on automation, visual coding and some long-horizon agent evaluations [24]. On DeepSWE v1.1, Xiaomi reports Pro moving from 58.4 to 72.57 and Flash from 48.8 to 65.68, gains of 14.17 and 16.88 points [10][23]. Those are Xiaomi's own evaluation, not independently reproduced [10]. For the deltas to transfer, your tasks have to be verifiable the way DeepSWE's are, and your graders have to agree with Xiaomi's about what a good trajectory looks like.

The MiMo team is led by Luo Fuli, the former DeepSeek researcher who joined Xiaomi to head it, VnExpress reported in November 2025 [20]. She published the final RL training runs live, internal metrics visible [21]. Xiaomi has not traditionally been counted among the six Chinese AI Tigers [26]. On what this means for closed APIs, @Yuchenj_UW claims frontier coding capability has plateaued since Opus 4.8 while open models keep closing the gap at 10 to 50 times lower cost [32]. @teortaxesTex argues the frontier has split into new higher tiers where internal and top closed models are still well ahead [33].

What to watch

  • An independent reproduction of the DeepSWE v1.1 numbers, which Xiaomi ran on its own evaluation.
  • Release of the complete 7,000-plus task datasets. Until they arrive, the published environments and recipes need someone else's tasks.
  • Whether the 46.32 Intelligence Index placing survives the next index version or the next open-weights release.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence60
Adoption25
Hype gap+30
Incentives65
Confidence62

Perspective Coverage

7 publishers
Builder
Builder 56%
Operator
Operator 23%
Investor
Investor 21%
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Xiaomi has released and open-sourced the MiMo-V2.6 series, two natively omnimodal models, MiMo-V2.6-Pro and MiMo-V2.6-Flash, for coding, visual tasks and computer use.

  2. [2]

    The MiMo-V2.6-Pro model card describes a sparse mixture-of-experts architecture with 1.02 trillion total parameters and 42 billion activated for each token.

  3. [3]

    The MiMo-V2.6-Flash model card lists 309 billion total parameters and 15 billion activated.

Sources

7 independent publishers whose own reporting we read for this story.

  1. blog.vercel.com

    1 article · September 20, 2026

    MiMo V2.6 models now available on AI Gateway
  2. huggingface.co

    1 article · September 22, 2026

    XiaomiMiMo/MiMo-V2.6-Pro-RL · Hugging Face
  3. latent.space

    1 article · September 21, 2026

    [AINews] Xiaomi MiMo-V2.6-Pro 1T-A42B: the new top Open Weights model, trained for $3M
  4. runtimewire.com

    1 article · September 21, 2026

    Xiaomi open-sources MiMo-V2.6 and the RL machinery behind it
  5. testingcatalog.com

    2 articles · September 21, 2026

    Xiaomi open-sources MiMo-V2.6 Pro and Flash models
  6. the-decoder.com

    2 articles · September 22, 2026

    Xiaomi's affordable flagship AI leads the open models, and Anthropic says Claude helped get it there
  7. thenewstack.io

    1 article · September 23, 2026

    “Impressive level of openness”: Xiaomi goes way beyond the usual open-weight playbook with MiMo-V2.6

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories