Skip to content

Build1 publisher3 min readPublished

Open 0.8B model Jeff answers agents' multiple-choice calls and escalates the unsure ones to a 27B

Jeff, an open 0.8B model, returns a probability for each option in one forward pass and sends low-confidence agent decisions to Qwen3.8-27B. The confidence score attached to each answer tells an agent when a small decision is worth the larger model's time.

The Engineer · Build desk

What happened

  • Nine LoRA adapters of about 41 MB each handle separate jobs, among them a prompt-injection guard, ticket triage, tool choice, grounding checks and spam.
  • The project's comparison ran both models on MLX on one M4 Max, with the 27B's step-by-step reasoning turned off.
  • On grounding, the 27B alone scored 96.7% and Jeff finished at 96.3%, a gap of one question in 300, while running 20 times faster.
  • In issue #7, an evaluation at a concurrency of two lost 100 of its 287 rows to the server rejecting overlapping requests.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint An agent that fans out tool calls has to write its own queue and retry logic before it can run the guard on every tool output.
  • decision Teams whose fallback is a hosted model, or has reasoning switched on, have to rerun calibration on their own rows before trusting the project's speed and accuracy figures.
  • cost Each decision run with orders: 2 pays for a second forward pass, so correcting position bias comes straight out of the latency budget.

The post's author argues that sending "which tool?" to a big model costs a prompt, seconds of waiting and a paragraph to parse, with no honest signal of how sure the model was [21]. Jeff's route is /v1/systemone on a small FastAPI server [10]. Inside, the model scores every option it was given, and a softmax turns those scores into probabilities [10]. It writes no text [2]. The usage block reports output_tokens: 0 on every call [11], the one token count in an agent stack that will never trip a budget alert. Each answer carries a key, a probability and a confidence running from 0, no better than a guess, to 1, certain [13]. In the README's client pattern, anything below a threshold goes to the big model [14], Qwen3.8-27B in the project's setup [3].

The base model takes any option list zero-shot [20]. Adapters cover the recurring jobs. According to the README, each one trained in a single epoch on one GPU, in half an hour to four hours [5]. All nine together come to roughly 370 MB [1].

Setting orders: 2 asks the question again with the options reversed and averages the two runs [12]. According to the post, models lean toward options near the top of a list, and the reversal cancels some of that [12]. I would turn it on for tool choice. The project's table puts a guard check at about a tenth of a second [16].

For the 20x speed on grounding [9] to transfer, a team's fallback has to match the configuration in the project's runs [7]. The post reports the cascade across eight adapters [6], one fewer than the nine that ship [4].

The calibration procedure is the part I would copy. Each threshold was fixed on its own calibration rows before any test row was scored [8]. Holding those rows apart means no threshold was tuned on the answers it was graded against. On the grounding row, where the small model trails, the post's author wrote that "A results table with a near-loss in it reads like a measurement, not a pitch." [22]

Adoption costs show up at concurrency. The server takes its lock without waiting. An overlapping request gets HTTP 529 with a one-second retry hint [17]. The loss in issue #7 works out to about 35 percent of the run [2]. The Python client retries nothing [19]. A guard placed in front of every tool output will hit that lock whenever two checks overlap.

For the five small calls the post traces through one support email [1], the evidence supports a cascade with Jeff first and the 27B behind it. It does not show the large model leaving the stack. The author adapted the client example from the README and wrote, "I did not run this" [15].

What to watch

  • A fix for issue #7 that makes the server queue overlapping requests or adds retries to the Python client.
  • Independent cascade runs against a hosted fallback model with reasoning on, testing whether the 20x speed and grounding figures hold.
  • Published cascade results for all nine adapters, up from the eight in the current comparison.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories