Build1 distinct publisher2 min readUpdated
Intern-S2-Preview-397B puts page-level paper reading and tool use into Apache 2.0 weights. The download listing says 810 GB, and every benchmark number so far is the lab's own.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Rendering a page and training on the picture is a claim about where scientific content actually lives. Shanghai AI Lab's visual pretraining works from rendered pages rather than text pulled out of PDFs, which keeps figures, tables, equations and layout attached to the prose around them [5]. The alternative is familiar to anyone who has pointed an extraction pipeline at a corpus of chemistry papers: table cells arrive in the wrong order and axis labels do not arrive at all. That decision, more than the parameter count, is what changes the class of question you can put to a paper the model has not seen.
The parameter count still deserves a look. The lab's name says 397B, while the Hugging Face listing says roughly 404 billion parameters and about 810 GB of files [11]. Four hundred and four billion values at two bytes each is 808 GB, so what is on offer is unquantized 16-bit weights rather than a packed release [19]. Label and listing differ by about 7 billion parameters, or 1.8 percent [22]. More usefully: 810 GB of weights spread over 80 GB accelerators occupies eleven of them before a byte of KV cache is allocated [20].
What the training pipeline is built for is harder than what the published scores measure. It combines supervised fine-tuning, reinforcement learning across scientific domains, reinforcement learning inside interactive agent environments and on-policy distillation, with a separate branch that forecasts time-series values numerically instead of spelling them out as text tokens [8][9]. The evaluation, meanwhile, capped text reasoning at 256,000 tokens and multimodal tasks at 64,000, and the authors say those limits describe the harness rather than evidence the model can finish an open-ended research programme without losing direction [7]. Serving that 256,000-token window sits on top of the eleven-device floor, not beside it. The breadth being attempted here follows chief scientist Bowen Zhou's stated thesis that a general model should keep broad capability while acquiring deeper expertise in individual fields [18].
The paper trailed the weights by about four weeks [21]. That ordering is now ordinary for open-weights releases, and the cost falls on whoever moved first: the earliest deployers ran a system whose training recipe and evaluation caveats were not yet written down [4]. Most teams will meet the model through the lab's OpenAI-compatible API in any case [12], which is the same arrangement they already have with closed models. Downloadable weights only change the relationship for the people who can hold them.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Shanghai AI Lab released the 397B model around July 16-18, 2026, according to the available research and the lab's API documentation.
The August 13 technical paper provided a fuller account of a model that was already available rather than marking a same-day launch.
Shanghai AI Lab reports competitive or leading results across its collection of scientific, multimodal, agent and time-series benchmarks, comparing the model with GPT-5-mini, Gemini 2.5 Flash, DeepSeek-V3 and its own larger Intern-S1-Pro.
Those results come from the Intern-S2 research team and have not been independently reproduced across the model's advertised range of tasks.
More than 120 researchers, among them Kai Chen and Wenwei Zhang, detailed Shanghai AI Lab's 397B-parameter bet on scientific agents in a technical paper published August 13.
Intern-S2-Preview-397B is designed to work across scientific documents, images, time-series data and external tools while sustaining the longer chains of activity required for research workflows.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Primary-document detail, single intermediary, no replication
The cluster's factual spine is unusually specific and traceable to primary artifacts: a dated technical paper, a model card, a Hugging Face listing with parameter and file-size figures, and API documentation. Arithmetic on the published 404B/810 GB figures is internally consistent and checkable. But every one of those artifacts reaches us through one publisher, no second outlet corroborates the mid-July release window, and all performance evidence is authored by the model's own research team. The only external check cited, SciAgentArena, addresses the general class of scientific agents rather than this model.
Distribution live, uptake undocumented
Real distribution exists: Apache 2.0 weights and deployment code on Hugging Face and InternLM GitHub, plus an OpenAI-compatible API, all live for roughly a month before the article. That is availability, not adoption. The supplied source names no external user, download count, derivative model, benchmark reproduction or production deployment, and it argues the 810 GB footprint, which needs at least eleven 80 GB accelerators, keeps smaller groups from meaningful experimentation. Scoring stays low because the only observed events are first-party release events.
Vendor framing runs ahead of the record; the coverage discounts it
The gap sits with the originating claims rather than the coverage. A 'scientific agentic foundation model' framed as research's 'next Coding' is asserted on the strength of the lab's own benchmark collection, with long-workflow reliability, verifiers and tool integration listed as unfinished, and evaluation token limits standing in for demonstrated long-horizon autonomy. That is overstatement relative to evidence and near-zero observed adoption. The score is only modestly positive because the single cluster source actively deflates it: it separates the paper from the launch, notes the paper is more careful than the model card, flags the parameter-label gap, and cites SciAgentArena's finding that agents remain uneven on open-ended research.
Model owner scores itself and sets the narrative
The party making the capability claims is the party that built, named, benchmarked and promotes the model. Shanghai AI Lab selected the benchmark collection and the comparison set, chose the 397B label, wrote the model card launch language, and its director frames scientific research as 'the next Coding', all of which serve an institutional positioning interest against commercial frontier labs. Countervailing signals exist and are why the score is not higher: Apache 2.0 weights and open deployment code invite scrutiny, and the paper voluntarily discloses remaining weaknesses. The publisher's own incentives are not characterizable from a single item.
Moderate: artifact-grounded but single-sourced
Confidence is middling. The structural facts, license, distribution channels, parameter and file-size listing, training pipeline components and the paper's own caveats, are specific and verifiable in principle, and the derived footprint arithmetic is sound. Against that, one publisher carries the entire cluster, the release date is an approximate window, and no independent evaluation, adoption or pricing evidence exists to test the capability and market claims. Assessment of the architecture and deployment burden is reasonably firm; assessment of performance and traction is not.
build
Armenian ASR leaderboard: closed models take the top eight, then lose the domains that matter1 distinct publisher
product
The cheapest model scored 10 out of 100: assistant choice is now a code-security decision1 distinct publisher
science
The real disclosure in Qwen3.8-Max is the rack: 2.4T open weights, 72 GPUs, 4K tokens/sec1 distinct publisher
product
Z.ai's GLM-5.3 beats Claude on CyberGym, then hands out the weights1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 21, 2026