Skip to content

Build1 publisher3 min readPublished

Ollama's new endpoint puts typed routing and moderation decisions on the local machine

Ollama 0.35 adds a /v1/systemone endpoint returning a choice, yes/no or score from models run on the device. Its 9B Nimble model matched human moderation labels 70.3% of the time in its maker's test, so each team still sets its own review threshold.

The Engineer · Build desk

Photograph accompanying Ollama's new endpoint puts typed routing and moderation decisions on the local machine
Photo: techcrunch.com

What happened

  • Ollama's September 29 release lists Bespoke Labs' Nimble and two experimental Together AI models, Tev1 at 4 billion parameters and a 0.8-billion-parameter version.
  • Nimble is fine-tuned from Qwen3.5-9B, scores answer tokens directly without a reasoning trace, and takes up to 64 questions about one input per call.
  • In Bespoke Labs' comparison, Nimble's macro-average accuracy was 74.8% against 76.0% for TypeSafe's hosted Jev, whose API shape the endpoint follows.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Choosing Nimble for moderation keeps content on the device at a cost in label agreement that the overall average understates by a wide margin.
  • constraint Teams cannot read a human-review cutoff off Nimble's confidence value, so each has to derive one from labelled items in its own queue.
  • cost Dropping per-request charges moves the bill to hardware, since every machine running Nimble must hold a 9.5 GB model or its 5.6 GB quantized build.

A call to /v1/systemone carries a `state`, which can be text or structured data, and a set of named questions [2]. Each question declares the kind of answer it wants: a choice, a yes-or-no, or a score [2]. Ollama's example sends one support ticket with three questions: which team should own it, whether it requests a refund, and how urgent it is [4]. The reply puts each answer in a defined field, with a probability for every allowed choice [3]. Routing code reads the field and branches. It does not parse generated text [3].

This is good engineering for narrow jobs. A dispatcher needs a closed answer set and a number it can compare, and the endpoint returns both [3]. With Nimble's limit of 64 questions per input [6], one call can carry a full triage checklist. The endpoint also follows the API shape of Jev, TypeSafe's hosted decision model [11].

The 91 ms average comes from a Pac-Man demonstration [8], and a ghost turning the wrong way costs less than a wrongly blocked customer. Ollama measured it on one MacBook Pro with an M5 Max [8]. For the figure to transfer, the target machine needs comparable memory and compute. The local route removes network transit and per-request hosted charges, and the hardware has to cover what is left [16].

The accuracy figures are Bespoke Labs' own, reproduced on Ollama's Nimble page [10]. They cover 3,880 examples across 13 public datasets [10]. That is about 298 examples per dataset on average [4]. Across all of them, Nimble trails Jev by 1.2 points of macro-average accuracy [2]. On Civil Comments, the moderation set, Jev agreed with human labels 81.0% of the time [12]. There the gap widens to 10.7 points [1]. Nimble disagreed with the annotators on 29.7% of those examples [3].

According to Runtimewire, these are agreement scores against dataset labels and do not prove performance on a company's own moderation queue [17]. For the Civil Comments number to hold, a company's content would have to resemble that dataset and its policy would have to match those annotators. I would not assume either for a support desk or a marketplace.

Nimble also returns a confidence measure. Ollama's documentation says it reflects how concentrated the model's choice is and does not promise the answer is correct [13]. A sharp distribution can still point at the wrong label. In my view the human-review cutoff has to come from labelled items in the real queue. I'd run a few hundred through Nimble, sort the errors by confidence, and draw the review line where the error rate stops being tolerable. Because the endpoint follows Jev's API shape [11], the same sample should run against the hosted model with little rework.

The launch post presents local use as carrying no additional inference cost, and Ollama has not published a separate price for the endpoint [14]. Ollama raised a $65 million Series B led by Theory Ventures in July, bringing its total funding to $88 million [15].

What to watch

  • Independent moderation results for Nimble or Tev1 on data outside Bespoke Labs' own 13-dataset comparison.
  • Accuracy figures for Together AI's experimental Tev1 models, and whether they leave experimental status.
  • Nimble latency on machines below an M5 Max, or for the 5.6 GB quantized build.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories