Skip to content

Build2 publishers3 min readPublished

llama.cpp brings TypeSafe's typed-decision format to five open model families on local hardware

llama.cpp merged a /v1/systemone endpoint on October 2 that returns typed answers with probabilities from five open model families running locally. Whether it can replace hosted classification calls depends on calibration that each team has to measure on its own data.

The Engineer · Build desk

What happened

  • TypeSafe AI introduced the /v1/systemone format with its hosted Jev model on September 15, as a way to return typed decisions to software workflows.
  • The llama.cpp pull request names Laya, Julia-1, Lev, OpenJev and Kev as the supported model families.
  • Xuan-Son Nguyen authored the change, and Georgi Gerganov reviewed and approved it and announced the endpoint on X.
  • Nguyen wrote on X that the feature should ship in llama.cpp v0.6.0 "early next week", with development builds usable until then.
  • OpenAI put its own Decisions API, built on its small Luna model, into limited preview, with broad release planned within days.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability Teams using a supported family can keep classification and routing requests on their own hardware without sending each one through a hosted service.
  • constraint A team whose preferred open model is outside the five families cannot use the endpoint until someone writes that model's conversion, metadata and server handling.
  • exposure Where a score decides whether a human looks at a moderation or escalation case, a miscalibrated probability sends cases past review, and only a team's own labeled cases can show it.

A call to /v1/systemone sends one shared state and several questions about it. In the pull request's example, the state holds a customer's message and order details [5]. The questions ask whether the customer wants a refund, which team should take the case, and how urgent or frustrating it is [5]. Each question declares an answer type: yes-or-no (`noul`), a choice among labels, or a score on an ordered scale [4]. The server returns a typed answer for each question with a probability distribution attached [5].

According to The New Stack, most teams handle this today by prompting a chat model to pick from a list and, with luck, reading token probabilities as a rough confidence score [16]. The other route is a small classifier. It is fast and cheap, but it needs labeled data and a retrain every time the label set changes [16]. A decision model takes new labels in the prompt and still returns a score a program can act on [17].

Support is per model. The pull request describes model-specific conversion, metadata and server handling, so loading some other open model does not give it the endpoint [19]. A community discussion had already sketched a wrapper built from existing llama-server features [8]. I think per-model conversion code belongs next to the server it has to track. The merged change puts it in the runtime, where conversion and server support are maintained alongside the project [8]. Nguyen wrote in the author disclosure that most of the code was AI-written and that he owned the design [10].

I'd treat the endpoint as ready to evaluate, but not yet as a proven replacement for hosted calls. The pull request compares llama.cpp outputs with reference outputs on example inputs [9]. A parity check is the right first test for a port. It shows the port reproduces the reference model on those cases [9]. It cannot show whether a 0.9 on the urgency question is right nine times in ten on a particular support queue [9]. Teams still have to run supported models on their own cases and check that reported confidence matches observed accuracy [18]. In practice that means holding out labeled real cases, grouping the endpoint's predictions by reported probability, and comparing each group's accuracy with its stated confidence.

OpenAI says its Luna-based decision model returns results in 150 milliseconds, against 1.6 seconds for GPT-6 Luna [12][13]. That works out to roughly 10.7 to 1 [1]. Both figures describe OpenAI's own models. They become a target for a local deployment only after a team times its own hardware on its own requests. The New Stack lists the per-call price, the number of candidate answers one request can carry, and whether developers can tune the model on their own data as still unclear [15]. An OpenAI spokesperson told the publication the company plans to share more "at broad rollout" [14].

What to watch

  • Whether llama.cpp v0.6.0 ships with /v1/systemone on the timeline Nguyen gave, or the feature stays on development builds.
  • Calibration results for any of the five families on real routing or moderation data, evidence beyond the pull request's parity checks.
  • OpenAI's per-call price and candidate-answer limit for the Decisions API at broad rollout, the figures a local-versus-hosted cost comparison needs.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories