Build3 distinct publishers3 min readPublished Updated
IBM's 3B, 8B and 30B dense models all get a thinking switch and native tool calling, but only the two larger ones get agentic RL, and the tuning mixture leans hard on software engineering.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The fine-tuning mixture repays a few minutes with a calculator. The agentic block comes to roughly 2.3 million samples [1], and software engineering accounts for 69 percent of it [7], which puts SWE trajectories at about 22 percent of everything the models were fine-tuned on [2]. The rest of that block is thinner: tool calling at 12.1 percent, terminal use at 8.0 percent, then math, search and action [7]. The listed slices only add to 93.6 percent, so part of the agentic corpus is unaccounted for in the post [9].
The non-agentic figures need a reading before they mean anything. IBM lists instruction following 18.8, coding 18.8, math 14.6, multilingual 7.0, science 5.4, reasoning 3.0 and safety 0.8 [8], and those add to exactly 68.4, the same number given for the non-agentic share of the whole corpus [3]. They are shares of the full mixture, then, not of their own block. Under that reading the category actually labelled reasoning is 3 percent of the mixture, about a seventh of the software engineering volume [4], while SWE trajectories plus non-agentic coding come to roughly 41 percent [5]. IBM describes 4.2 as the release that adds explicit reasoning to a family that was previously instruction-following assistants [12]. By data volume it is a code and terminal model that can also think out loud.
Two more numbers from the same paragraph of the build post. About 35 of the 100 billion SFT tokens are not trainable [6], so roughly a third of that corpus is context the gradient never touches. And the entire supervised stage is under 0.7 percent of the pre-training token budget [7], which is a thin coat over a conventionally trained base.
On serving, the attention layout is 40 query heads to 8 KV heads [9], a five to one ratio [8]. That ratio is what makes the long-context phase [3] tractable to serve: cache is one fifth of what an equivalent multi-head configuration would want at the same sequence length. RoPE theta is set to 10 million, weights are bfloat16, input and output embeddings are untied [9]. Nothing unusual in the stack, which is rather the point, because at long sequence lengths the capacity question is cache rather than parameters.
The agentic trajectories were generated across a named set of scaffolds, among them OpenHands, SWE-agent, Terminus-2, Codex and Goose, mixing open datasets with IBM's own synthetic RL environments [11]. That breadth is also a distribution: a model trained to call tools inside those harnesses has learned their conventions. IBM's integration claim is that an OpenAI-compatible endpoint under vLLM emits function calls in OpenAI format and plugs into agentic harnesses without extra glue, with SGLang supported as well [10]. That is a claim to test against your own harness rather than accept.
What the material does not contain is a scoreboard. The post walks through architecture, data and the RL stages, and there are no evaluation figures here to place the 30B against anything else. Apache 2.0 across all three sizes [2] is what earns the family an evaluation slot at all; the mixture is what tells you which evaluation to run first.
Ranked by verification strength, evidence, and original report placement.
Granite 4.2 is IBM's first family of dense, decoder-only reasoning LLMs, released in three sizes: 3B, 8B and 30B, all sharing the same architectural design and training pipeline.
All Granite 4.2 models are released under the Apache 2.0 license.
Each model is pre-trained from scratch on approximately 15 trillion tokens using a five-phase strategy; phase 5 introduces long-context training, extending the context window to 512K tokens.
Every Granite 4.2 model has a thinking / non-thinking switch, a low-effort thinking mode that spends a short reasoning budget on easy questions, and native tool calling.
The 8B and 30B models additionally go through an agentic RL block that teaches them to call tools, edit and run code, drive a terminal and search the web inside real sandboxed environments.
The SFT data mixture combines agentic (31.6%) and non-agentic (68.4%) data, totalling approximately 7.2 million samples, or roughly 100B tokens, of which about 65B are trainable.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed first-party disclosure, no external verification
The cluster rests on one source, but that source is unusually concrete: exact architecture parameters, token budgets, sample counts and mixture percentages, all internally consistent under arithmetic checks. What is missing is any evaluation result (the results section is truncated in the supplied body) and any independent confirmation, which caps evidence well below high.
Fresh release with runtime support, no usage evidence
Adoption signal is limited to the release event itself plus vendor-stated support in vLLM and SGLang. There are no deployments, downloads, third-party benchmarks or customer usage disclosures in the supplied material, so the score reflects availability rather than uptake.
Mildly overstated: capability framing outruns shown results
The build disclosure is sober and the story's own reading of the mixture is arithmetically grounded, so the gap is small. It is positive because 'reasoning' and 'agentic' capability framing carries the release while no evaluation numbers appear in the supplied text, and because the mixture itself is dominated by software engineering data (~41% with coding) against 3.0% labelled reasoning.
Vendor-authored launch post on a distribution platform
The single source is written by IBM's Granite Team and published on Hugging Face, the platform where the weights are distributed. The publisher and the subject are the same party, and the post's purpose is to drive adoption of an open-weight release, so promotional incentive is high even though the technical content is verifiable.
Moderate: specific but single-sourced and partly truncated
Confidence is moderate because the factual spine is precise and self-consistent, but every claim traces to one vendor source, the captured text ends mid-section before results, and two published breakdowns are ambiguous or incomplete. Structural and licensing facts are reliable; capability quality is not yet assessable.
build
Intel puts its Arc GPU operating knowledge inside the coding agent already installed1 distinct publisher
build
Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters1 distinct publisher
build
Meta's real announcement is the split: 30B on your GPU, everything else behind the API6 distinct publishers
build
Your vLLM Manifest Would Boot SGLang Too, And That Is the Problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 25, 2026
1 article · August 26, 2026
1 article · August 25, 2026