Build1 publisher3 min readPublished
A 30B model with 3B active arrives on JumpStart, aimed at the cheap middle of agent work
NVIDIA's Nemotron 3.5 Lightning is now deployable from SageMaker JumpStart and fits on one GPU. The argument underneath it is about billing, not benchmarks.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- NVIDIA Nemotron 3.5 Lightning is now available in Amazon SageMaker JumpStart.
- With this launch, users can deploy Nemotron 3.5 Lightning from Amazon SageMaker JumpStart without configuring the serving infrastructure themselves.
- At 30B total parameters with only 3B active, Nemotron 3.5 Lightning can run on a single supported GPU.
- Its Mixture-of-Experts architecture activates 3B of 30B parameters per forward pass, helping maintain high throughput across long, multi-turn sessions.
- Nemotron 3.5 Lightning is a publicly available foundation model distilled from NVIDIA's frontier Nemotron 3 Ultra and developed with the Nemotron Coalition.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
NVIDIA Nemotron 3.5 Lightning is now available in Amazon SageMaker JumpStart, which means it can be deployed without configuring the serving infrastructure yourself [1][2]. What makes it worth reading is the shape of the model: 30B total parameters with only 3B active, which NVIDIA says lets it run on a single supported GPU [3].
That ratio is the whole pitch. The hybrid Mixture-of-Experts architecture activates 3B of 30B parameters per forward pass, roughly 10 percent, which NVIDIA says helps hold throughput across long multi-turn sessions [4][10]. The model is distilled from NVIDIA's frontier Nemotron 3 Ultra and developed with the Nemotron Coalition, trained specifically for agentic tool use across popular agent harnesses [5][6]. It is trained on open datasets and released as an open model, so you can customize it, own the resulting weights, and deploy it where your agents already run [7].
The reasoning in the announcement is more interesting than the marketing line about it being the fastest open model in its class for always-on agents [8]. NVIDIA's own framing is that agent steps are not fungible: planning a multi-stage workflow or orchestrating sub-agents can demand frontier-level reasoning, while classifying an alert, extracting fields from a form, or checking a record against a policy often does not [12][13]. Those cheap steps, per the post, can account for a large share of total call volume [14]. Routing all of them through one large model adds frontier-model cost and latency to work a smaller specialized model can do, so the proposed alternative is a system of models where each step goes to a model suited to it [15]. If you run NVIDIA NeMo Switchyard, it can do that routing across your chosen model pool [18].
Two supporting details matter operationally. DFlash speculative decoding can further reduce per-token latency [16], and a 1M-token context window lets an agent carry accumulated state through a long session without repeated re-grounding [17]. NVIDIA claims up to 4x higher throughput and up to 30 percent faster task completion on high-volume agentic workloads [9], but the post states those as "up to" figures without naming the comparison baseline [25]. Treat them as directional until you measure your own step mix.
On accuracy, the post says NVFP4 stays close to BF16 on many tasks across the published evaluations [19], with recipes and commands published in NeMo Gym [20]. NVIDIA also states that the numbers were measured under its own consistent harness and may differ from vendors' self-reported results [21], which is a useful admission and a reason to rerun the harness yourself.
One gap to plan around: NVIDIA says organizations can post-train the model with NeMo for domain-specific tools, workflows, and policies, then deploy the result in their chosen environment [22], but the SageMaker JumpStart model card for this launch does not expose JumpStart customization [23]. So the fine-tuning happens elsewhere and the serving happens here.
The listed target workloads are the high-frequency ones: personal assistants over email and calendar, document extraction and policy checks in financial services, alert enrichment and incident classification in security operations, and network alarm triage in telecom [24].