Skip to content

Build1 publisher2 min readPublished

Nex-AGI's runnable N2.5 tiers ask for two H100s or sixteen H200s

Publishing weights and publishing something a team can actually deploy are different acts, and N2.5 lands on both sides of that line depending on which tier you pick and whether its weights exist yet.

The Engineer · Build desk

Illustration accompanying Nex-AGI's runnable N2.5 tiers ask for two H100s or sixteen H200s

What happened

  • Nex-AGI released N2.5 as three agent tiers: a 35B Mini and a 397B Pro built for operating computers and browsers, plus a 1.6T Max that drops visual input for reasoning, coding, scientific work and agent workflows.
  • Nex's sample deployment puts Mini on two H100s, while the reference setup for Max calls for sixteen H200s spread across two nodes.
  • The repository lists 3B active parameters for Mini, 17B for Pro and 49B for Max, because the mixture-of-experts design fires only part of each model for every token.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision A team wanting multimodal computer use on open weights now picks between a 3B-active model that fits two GPUs and a 17B-active one it cannot obtain at any price.
  • cost There is no single capacity plan for "running N2.5": choosing Max commits you to multi-node serving before you have evaluated a single trajectory.
  • constraint The largest downloadable tier cannot take an image, so interface-feedback work is capped by whatever Mini can do until Pro ships.
  • contradiction The strongest computer-use score in the announcement sits on weights nobody outside Nex can load, so the people who would deploy it are the ones who cannot check it.

Sixteen H200s across two nodes is eight per node [17]. That split is a fabric requirement as much as a GPU count. Max holds 1.6T total parameters and activates 49B per token [2][6], roughly 3.1 percent of the model [16]. Cheap tokens do not shrink the weights you have to keep resident, and residency is what pushes Max's reference configuration to eight times the GPU count of the two-H100 sample deployment Mini runs in [18][4]. Mini itself is 35B with 3B active [2][6], about 8.6 percent [20].

The tier in between is where the release's evidence sits. Pro is 397B with 17B active [2][6], it takes images as well as text [3], and Nex reports 56.4 on OSWorld-2 for it against 46.7 for Qwen3.8-Max [9][10], a margin of 9.7 points [19]. Nex also says Pro drove more than 468 actions in Pokemon Platinum using visual feedback and standard controls [11]. On September 8, 2026 its model card was still marked coming soon, with no downloadable weight shards [5]. That leaves Mini and Max as the checkpoints a team can pull, and Max takes no visual input at all [3].

The AutomationBench v1.0.6 and OSWorld-2 figures are Nex's own, and the reviewed materials do not replicate them independently [9]. For the 9.7-point margin to say anything about your automation you would need the harness, the step budget and the retry allowance to match; Nex has not published enough methodology in those materials to establish direct comparability with the Qwen3.8-Max number [10]. Until that arrives, 56.4 is a statement about Nex's scaffold on Nex's task set.

Adoption cost has a floor. It is lower than the GPU counts suggest. Mini and Pro are post-trained from Qwen3.5 variants, and Max is based on DeepSeek-V4-Pro-Base [8]. A serving stack that already loads those base architectures treats N2.5 as weights and config rather than a kernel project. Nex's earlier open-source work shipped datasets, agent frameworks, reinforcement-learning tooling and inference infrastructure alongside models [13], which is the layer that decides whether anyone runs a checkpoint a second time. That is genuine systems work. It is the reason the two-H100 path is credible at all.

Counterparty is the odd part. Nex-AGI's public materials describe a collaboration involving the Shanghai Innovation Institute, Shanghai Qiji Zhifeng, Mosi Intelligence and Kuafu Technology, with no conventional standalone startup, no named founder or chief executive, and no publicly established legal structure [12]. The N1 technical paper is dated December 4, 2025 and credits a large research team [14]. It is a setup well suited to running research, but not one that maps easily onto a procurement form.

What to watch

  • Pro weight shards appearing in the repository, which would make the 56.4 OSWorld-2 claim testable outside Nex.
  • A single-node reference configuration for Max, which would move the planning number off sixteen H200s.
  • Published harness and step-budget details for the AutomationBench v1.0.6 and OSWorld-2 runs.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories