Skip to content

Build1 publisher2 min readPublished

OpenAI and AWS reportedly answer Jev with decision models of their own within two weeks

OpenAI and AWS shipped their own decision models within two weeks of TypeSafe AI's September 15 launch of Jev, according to a dev.to account. For routing work, the comparison with an LLM call turns on whether a decision model's probabilities are calibrated.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying OpenAI and AWS reportedly answer Jev with decision models of their own within two weeks
Generated illustration

What happened

  • Jev takes program state, usually text or JSON, plus typed questions (yes or no, a pick from a fixed list, or a score against a rubric) and returns only probabilities.
  • TypeSafe AI's founder, Diogo Almeida, contributed to OpenAI's instruction-following work, the effort that turned into ChatGPT.
  • OpenAI announced a Decisions API built on its small Luna model ahead of DevDay, and AWS shipped Strands Decider 2B, a model small enough to run locally.
  • A 330-comment r/LocalLLaMA thread called Jev old tech in a new costume, arguing that calibrated classifiers are a solved problem.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Because output is free, spend tracks the state sent in: a million replies scored at the SillyTavern rate would cost about $500.
  • constraint Jev's thin niche-domain knowledge puts the context on the caller, so it suits routing and triage better than calls that need expert judgement.
  • decision Teams classifying with LLM calls now have a probability-returning option to benchmark, and the deciding measurement is whether confidence tracks accuracy on their own labelled traffic.
  • precedent With two large platforms shipping versions within two weeks, the decision-model pattern no longer depends on one startup's pricing or survival.

A generative model emits one token at a time, each conditioned on the last [8]. Jev skips that loop. According to the dev.to post, it samples every answer at once, in parallel and hardware aware, and the post credits that step for the speed [8]. All the typed questions in a request come back from a single call [6]. The whole interface sits behind one endpoint, api.typesafe.ai/v1/systemone [7].

Speed is the smaller claim. TypeSafe trains Jev with RLCD, Reinforcement Learning for Calibrated Decisions, aimed at honest probabilities where high confidence means high accuracy; RLHF instead optimizes for what human raters prefer [9]. I think calibration is what decides whether a router can trust the model to hand off the hard cases. "If a model can do a task 95 percent of the time but cannot say when it is in the failing 5 percent, you cannot wire it into anything," the post's author wrote [10]. The same post says TypeSafe has not published enough architecture detail to verify the interesting claims [11].

Input costs $0.042 per million tokens, and output costs nothing [5]. According to the post, the company admits it cannot prove the pricing is not subsidized [12]. "That price is a bet, not a fact," the author wrote [13].

A model pitched for industrial automation drew its busiest early threads from r/SillyTavernAI, a roleplay tool, and r/SideProject, where people scored startup ideas [19]. The SillyTavern users report scoring replies in about 200 milliseconds for roughly $0.0005 a call [20]. For one verdict on one reply, that is a lot of input. If the whole charge is input at the list rate, $0.0005 divided by $0.042 per million tokens comes to about 11,900 tokens per call [1].

TypeSafe quotes 70 to 500 milliseconds per request [4], and the post puts OpenAI's Luna-based API at about 150 milliseconds per call [17]. Both are measurements of someone else's payload. I'd expect latency to move with how much state goes in, and the SillyTavern figure shows state can run past 10,000 tokens [1].

On this record, the OpenAI and AWS releases are single-sourced to the dev.to post. It quotes a TechCrunch headline: "Amazon releases its own Jev clone as decision models flood the web." [18] The post does not give prices for either release. Free output is documented only for Jev [5].

What to watch

  • Published pricing for OpenAI's Decisions API and AWS's Strands Decider 2B, and whether either matches Jev's free output tokens.
  • Architecture or calibration data from TypeSafe that lets outsiders check the RLCD claims on independent workloads.
  • OpenAI's DevDay documentation for the Decisions API, showing whether it uses the same yes-or-no, choice and score question types.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories