Build1 publisher3 min readPublished
Ollama's new endpoint puts typed routing and moderation decisions on the local machine
Ollama 0.35 adds a /v1/systemone endpoint returning a choice, yes/no or score from models run on the device. Its 9B Nimble model matched human moderation labels 70.3% of the time in its maker's test, so each team still sets its own review threshold.
The Engineer · Build desk

What happened
- Ollama's September 29 release lists Bespoke Labs' Nimble and two experimental Together AI models, Tev1 at 4 billion parameters and a 0.8-billion-parameter version.
- Nimble is fine-tuned from Qwen3.5-9B, scores answer tokens directly without a reasoning trace, and takes up to 64 questions about one input per call.
- In Bespoke Labs' comparison, Nimble's macro-average accuracy was 74.8% against 76.0% for TypeSafe's hosted Jev, whose API shape the endpoint follows.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Choosing Nimble for moderation keeps content on the device at a cost in label agreement that the overall average understates by a wide margin.
- constraint Teams cannot read a human-review cutoff off Nimble's confidence value, so each has to derive one from labelled items in its own queue.
- cost Dropping per-request charges moves the bill to hardware, since every machine running Nimble must hold a 9.5 GB model or its 5.6 GB quantized build.
A call to /v1/systemone carries a `state`, which can be text or structured data, and a set of named questions [2]. Each question declares the kind of answer it wants: a choice, a yes-or-no, or a score [2]. Ollama's example sends one support ticket with three questions: which team should own it, whether it requests a refund, and how urgent it is [4]. The reply puts each answer in a defined field, with a probability for every allowed choice [3]. Routing code reads the field and branches. It does not parse generated text [3].
This is good engineering for narrow jobs. A dispatcher needs a closed answer set and a number it can compare, and the endpoint returns both [3]. With Nimble's limit of 64 questions per input [6], one call can carry a full triage checklist. The endpoint also follows the API shape of Jev, TypeSafe's hosted decision model [11].
The 91 ms average comes from a Pac-Man demonstration [8], and a ghost turning the wrong way costs less than a wrongly blocked customer. Ollama measured it on one MacBook Pro with an M5 Max [8]. For the figure to transfer, the target machine needs comparable memory and compute. The local route removes network transit and per-request hosted charges, and the hardware has to cover what is left [16].
The accuracy figures are Bespoke Labs' own, reproduced on Ollama's Nimble page [10]. They cover 3,880 examples across 13 public datasets [10]. That is about 298 examples per dataset on average [4]. Across all of them, Nimble trails Jev by 1.2 points of macro-average accuracy [2]. On Civil Comments, the moderation set, Jev agreed with human labels 81.0% of the time [12]. There the gap widens to 10.7 points [1]. Nimble disagreed with the annotators on 29.7% of those examples [3].
According to Runtimewire, these are agreement scores against dataset labels and do not prove performance on a company's own moderation queue [17]. For the Civil Comments number to hold, a company's content would have to resemble that dataset and its policy would have to match those annotators. I would not assume either for a support desk or a marketplace.
Nimble also returns a confidence measure. Ollama's documentation says it reflects how concentrated the model's choice is and does not promise the answer is correct [13]. A sharp distribution can still point at the wrong label. In my view the human-review cutoff has to come from labelled items in the real queue. I'd run a few hundred through Nimble, sort the errors by confidence, and draw the review line where the error rate stops being tolerable. Because the endpoint follows Jev's API shape [11], the same sample should run against the hosted model with little rework.
The launch post presents local use as carrying no additional inference cost, and Ollama has not published a separate price for the endpoint [14]. Ollama raised a $65 million Series B led by Theory Ventures in July, bringing its total funding to $88 million [15].
What to watch
- Independent moderation results for Nimble or Tev1 on data outside Bespoke Labs' own 13-dataset comparison.
- Accuracy figures for Together AI's experimental Tev1 models, and whether they leave experimental status.
- Nimble latency on machines below an M5 Max, or for the 5.6 GB quantized build.