Skip to content

Product1 publisher3 min readPublished

Xiaomi's MiMo-V2.6-Pro leads the open-weight index at $0.87 per million output tokens

Artificial Analysis scores the new Xiaomi model first among open weights and twelve points behind Claude Opus 5.5, and the cheaper Flash tier is the one an operator should put in front of a real queue.

The Product Desk · Product desk

Illustration accompanying Xiaomi's MiMo-V2.6-Pro leads the open-weight index at $0.87 per million output tokens

What happened

  • Xiaomi released two open-weight models, MiMo-V2.6 Pro and Flash, under an MIT licence on Tuesday, with the weights posted to Hugging Face and both models live on its own API and on OpenRouter.
  • Artificial Analysis scores Pro at 46 on its Intelligence Index, first of the 114 models in its class and the highest any open-weight model has reached, and also rates it the cheapest model it tracks at $0.13 per task.
  • Pro lists at $0.435 per million input tokens and $0.87 per million output, and Anthropic cut Claude Opus 5.5 to $4 and $20 on the same day.
  • Xiaomi says both models came out of a single reinforcement learning run of 30 steps over roughly 750,000 trajectories in under six days, costing about $2.62m for Pro and $0.85m for Flash.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Teams that ruled out open models on quality in an architecture doc owe that line a re-test, and the cheapest version of the test is a hosted endpoint.
  • cost Keeping a frontier model on bounded, output-heavy work means paying the premium on every generated token each month, whether or not the task needed the extra index points.
  • constraint The remaining quality gap sits in long autonomous terminal sessions, so a frontier vendor stays in the stack for that queue even if bulk agent traffic moves off it.
  • capability A small team that has never built an evaluation harness can now point published graded environments at its own agents, independent of which model it ends up buying.

Put the two output prices side by side and the ratio is 23 to 1 [25]. On input it is closer to 9 to 1 [26]. For a team running agent loops that write long, that ratio shows up on the monthly invoice, not in the index score.

What is being sold here is an open-weight model at the top of an independent index. Xiaomi opened its announcement with "Frontier intelligence, all the modalities, built in public" [2]. What most teams will do is call it over OpenRouter and never touch a GPU [9]. TNW's account of the release did not report what hardware it takes to serve 1.02 trillion parameters, 42 billion of them active per token [5][34]. Capability was rarely the only reason self-hosting got ruled out.

The capability case is better than it was and still not level. Artificial Analysis has Claude Opus 5.5 at 58, twelve points above MiMo [18][28]. Xiaomi's own tables put Claude Opus 5 ahead of Pro on DeepSWE and on ProgramBench, and 14.1 points ahead on Terminal Bench 4.0, where GPT-6 Astra scores 59.6 [17][29]. Xiaomi also says Pro performs on par with Opus 5 and GPT-5.6 Sol across most agent benchmarks; that one is the company's own testing [19].

Flash is the tier to test first. It costs $0.14 and $0.28 per million tokens, roughly a third of Pro, and keeps the million-token context and the multimodal input [14]. Against Opus 5.5 on output that is a factor of 71 [27]. On Xiaomi's agent benchmarks it trails Pro by four points or fewer, at 67.9 against 71.9 on DeepSWE and 61.2 against 62.0 on JobBench [15]. On CyberGym it beats Pro, 95.1 to 94.0 [16].

Xiaomi reports DeepSWE twice, and the two figures do not match. The training run ends at 72.57 for Pro and 65.68 for Flash [13]. The benchmark table lists 71.9 and 67.9 [15]. That is 0.67 of movement for Pro and 2.22 for Flash [30]. A team choosing between the tiers on a four-point DeepSWE gap is working inside the spread of the vendor's own numbers.

The swap turns on whether the work is a long autonomous terminal session or a bounded task you can write a grader for, and on whether output tokens dominate the bill. If the work is output-heavy and bounded, Flash goes on a shadow queue behind the current model and you compare completions on real tickets. If it is long-horizon terminal work, the 14.1 points on Terminal Bench 4.0 is what the lower price costs, and that number comes from Xiaomi's own table [17][29]. For latency-bound work there is Pro-UltraSpeed, which Xiaomi says runs up to 20 times faster than Pro at the same quality for ten times the token price, $4.35 in and $8.70 out [10].

For a team that has never written an automatic grader, the reusable part of this release is the harness. Xiaomi published the technical report, the end-to-end reinforcement learning framework, more than 7,000 task environments with automatic graders, a set of composable mini-harnesses and a smaller distilled model [20]. The environments cover software development, cybersecurity, office work and web design, with roughly a thousand more on music composition [21]. Xiaomi also published the reward design, the hyperparameters, the data mixtures and the costs, saying the point is reproducibility [22]. One detail is worth copying into your own eval setup: early agents solved assigned bugs by fetching the fix from a later version of the package, so Xiaomi stripped build caches and future Git history from the training environments and ran a dedicated agent to hunt for remaining loopholes [24].

What to watch

  • Whether the 46 holds once Moonshot's Kimi K3 and Alibaba's Qwen3.8 Max answer, since those two have traded the open-weight lead for most of this year.
  • Whether an outside lab reproduces the six-day run from the published framework, hyperparameters and data mixtures.
  • Whether anyone publishes serving requirements and real costs for hosting 1.02 trillion parameters, the figure a self-hosting decision turns on.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories