Skip to content

Build1 publisherNot yet confirmed elsewhere2 min readPublished

Ollama and llama.cpp now serve Jev-spec decision models on local hardware

Ollama and llama.cpp added local endpoints within a week for Jev decision models, 144M-to-27B checkpoints that return a label and a probability. Routing calls can move off hosted models once teams check each checkpoint's licence and test it on their own labels.

The Engineer · Build desk

How we use AISend a correction

What happened

  • Jev is an API contract defined by TypeSafe: a model takes a question and a fixed list of options and returns a probability for each one.
  • Ollama v0.35 shipped its /v1/systemone endpoint on September 29, with the Nimble and Tev1 checkpoints available to pull at launch.
  • llama.cpp 0.6.0 followed on October 5, serving five GGUFs (Julia-1, Laya, Kev-4B, lev and OpenJev), and Nimble was added days later.
  • OpenJev is licensed CC BY-NC, which bars commercial use, while the other five llama.cpp checkpoints ship under Apache-2.0.
  • Bespoke Labs' comparison of 3,880 human-labelled decisions scored local Nimble at 75.7%, against 76.0% for hosted Jev 1.13.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability Jobs that could not justify a hosted-model call on every request, such as tagging every incoming ticket, become cheap enough to run on a small local checkpoint.
  • decision Picking a llama.cpp checkpoint means picking mostly unbenchmarked weights, since Nimble is the only one of the six with a published accuracy figure.
  • exposure A commercial pipeline that pulls OpenJev takes on a non-commercial licence, and llama.cpp's MIT runtime licence does not clear it.
  • constraint Until someone measures calibration, the allow/deny threshold for these models has to be set from each team's own labelled data.

The best engineering here is the response shape. A routing call comes back as `{ "route": "billing", p: 0.87 }`, in one pass, with the same shape on every call [13]. Ask a chat model for the same decision and you get a sentence worded differently each time, plus a parsing step that fails quietly, according to the explainer that documented the release [14].

The checkpoints run from 144M to 27B parameters [6]. According to the explainer, the 144M to 421M pair runs on an idle CPU and handles clean routing. The 4B pair is for subtler labels, and the 27B adds vision [19]. Those descriptions come from the Ollama announcement, the llama.cpp 0.6.0 notes and the checkpoint cards [11]. The post is from mrsaynothing.dev, an agent-run site that publishes one post a day [12]. Its author says the models were read about, not run [11]. The site tells readers to benchmark on their own labels before it has run a bench of its own, and it says so; measured latencies are promised in a follow-up [10][11].

Bespoke Labs' table also scores two Tev1 sizes: 73.3% at 4B and 63.5% at 0.8B [9]. Dropping from 4B to 0.8B costs Tev1 9.8 points [16]. The test spans 13 datasets, about 298 decisions each on average [9][18]. The explainer calls the figures vendor-adjacent [10]. For any of them to carry over, a team's traffic would need roughly the class balance and ambiguity of those 13 datasets.

Accuracy is the right figure for a router, because a router takes the top option. A gate allows or denies on a threshold. That means it needs calibration, which the explainer defines this way: a calibrated 0.9 is right 9 times in 10 [15]. The post does not report calibration for any checkpoint. I'd move clean routing first. A miscalibrated 0.87 still picks the right queue as long as the ranking holds [13]. Gates can stay on their current model until the probabilities have been checked against local labels.

The licence check has to be done per file and per catalogue. Ollama's library carries its own set, Nimble, Tev1, Laya and Clef, behind its own /v1/systemone [3][1]. The explainer's account of the llama.cpp shelf is also inconsistent. It counts six checkpoints, including Nimble, which arrived days after the release. Then it describes a seventh, Clef, as shipping in the same 0.6.0 release under Apache-2.0 for text and vision [2][4].

What to watch

  • The follow-up bench promised by mrsaynothing.dev, which would put measured latencies against the size ladder the checkpoint cards describe.
  • Calibration results for any Jev checkpoint on independently labelled data, the figure that gating thresholds depend on.
  • Accuracy numbers for Julia-1, Laya, Kev-4B, lev and OpenJev, none of which appear in the Bespoke Labs comparison.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories