Skip to content

Product2 publishers3 min readPublished

AWS open-sources Strands Decider 2B to make agent routing decisions on a laptop

AWS open-sourced Strands Decider 2B, a model that picks from a fixed list of options in under 150 milliseconds on a local machine, the company says. Agent builders get a small model they can run themselves for routing and tool-selection steps that would otherwise each call a full LLM.

The Product Desk · Product desk

Illustration accompanying AWS open-sources Strands Decider 2B to make agent routing decisions on a laptop

What happened

  • AWS built the model on Qwen3.5-2B's base and swapped its text-generating head for a pointer head of roughly 1 million parameters that scores the listed options.
  • AWS aims it at agent steps including model routing, tool selection, context management, guardrail enforcement and policy classification.
  • TypeSafe AI's Jev drew attention to decision models a couple of weeks earlier, and AWS's team says Jev's parallel output design struggles with intricate reasoning.
  • AWS reports high accuracy and calibration on JevBench when its model is compared with other open-source 2B models.
  • The weights are on Hugging Face, with the full code, training scripts and examples published on GitHub.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost A first trial costs a team the time to label examples from its own routing logs and run them locally, with no per-call API bill to approve beforehand.
  • constraint Any step whose outcome gets challenged, such as a guardrail that blocks a customer, needs its reasons recorded by another system, because the decider cannot explain itself.
  • decision Because the decider runs on a laptop or in a public cloud, teams have to choose where it sits, and placing it beside the agent takes a remote model call out of every branch.

Take an agent step whose only job is to choose which of four tools should handle an incoming support ticket. On a general LLM, that step means the model writes a short answer, spending tokens and time to do it [2], and the code then parses the answer back into one of the four options. Strands Decider 2B is handed the options and returns one of them, with no other output [3]. Each pick arrives with a confidence score meant to show how likely it is to be right [4].

AWS pitches the release as a way to speed up agentic AI development more broadly [1][17]. What it ships is narrower and easier to test. The head AWS added is about 0.05 percent of the model's parameters [1]. The rest is an existing open LLM, tuned with a rank-16 LoRA adapter so the head can score the listed choices directly against answer positions [9]. AWS has labeled this release v.20 after repeated rounds of iteration [10]. Put plainly, it is a small classifier on top of a small LLM, and a team can check it against labeled examples from its own logs.

So far the evidence is AWS's own, on a benchmark that carries the name of the rival model its team criticises [11][12][13]. The coverage does not include the scores, or any test against a large LLM making the same calls.

AWS's own design for using it is a hybrid agent: the decider takes simple, repetitive choices and an LLM takes complex reasoning [15]. In that design the confidence score decides the handoff. A pick above a set threshold stands, and anything below it goes to the larger model. That split saves money only if high-confidence picks are right about as often as the score says, so for routing the calibration half of AWS's JevBench claim matters more than the accuracy half [11]. If most calls land below the threshold, the agent pays for both models.

Two axes sort the candidate steps. One is whether every option can be written down before the call, since the decider only chooses among choices it is handed [3]. The other is what a wrong pick costs. Fixed options with a cheap mistake, such as a tool call that can be retried, are where a trial makes sense now. Fixed options with an expensive mistake, such as a guardrail that lets the wrong request through, can still use the decider behind a high threshold, so doubtful calls go to the LLM. Steps with open-ended output stay on the LLM whatever a mistake costs. For the steps that qualify, the evidence a team needs is its own fallback rate and error rate on real traffic, set against what the step cost when it ran on the LLM.

What to watch

  • AWS or an independent tester publishing JevBench scores, or a head-to-head against a large LLM on the same routing and guardrail tasks.
  • Whether AWS builds Strands Decider into its managed agent services or leaves it as a Strands Labs download.
  • A Jev update from TypeSafe AI that answers AWS's claim about its performance on intricate reasoning.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories