Skip to content

Build2 publishers3 min readPublished

Typesafe's Jev API gains an open-weight rival in Cloudflare's Clef decision models

Cloudflare released two Jev-API-compatible decision models, Clef and Clef-flash, on Workers AI and as Apache 2.0 weights on Hugging Face. Typed classification steps in agent code can now move between providers or onto owned hardware, as long as they stay inside the text-only, 32k-context features Jev supports.

The Engineer · Build desk

Illustration accompanying Typesafe's Jev API gains an open-weight rival in Cloudflare's Clef decision models

What happened

  • In a Cloudflare domain-classification test, Clef took 2.2 seconds to fetch, render and classify a site, while gpt-oss-120b took 4.7 seconds and returned two labels.
  • On Typesafe's own eval suite, run by Cloudflare, Clef beat Jev in three of four areas, according to Cloudflare.
  • Cloudflare also launched a reinforcement-learning product that lets customers fine-tune Clef for their own use cases.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability Decision steps coded to Jev's API can switch between Typesafe and Workers AI without changes to the code that consumes the typed answers.
  • constraint Any step that sends images or more than 32k of context can fall back only to Clef, hosted or self-run, and has no fallback at Typesafe.
  • cost Running the Apache 2.0 weights locally puts the serving layer, including matching Jev's request and response shapes, on the team that runs them.
  • decision Teams already on Jev have to time Clef against Jev on their own inputs before switching, because the published timing compares Clef with a general LLM.

A decision call takes an input and a question, for example whether a support message is urgent and which team should handle it [8]. It returns typed answers with probabilities, and the calling code uses them to route the ticket, escalate it, or defer to a human [8]. Cloudflare says the model handles new categories without retraining [2]. Its case for the category is consistency. It describes LLMs as largely non-deterministic [15] and decision models as producing bounded structured outputs cheaply, quickly and consistently [1].

The probabilities are per label. In the Threat Intelligence team's domain test, Clef scored one site at 95% fashion, 85% ecommerce and under 1% phishing [9]. The first two sum to 180%, so the labels are scored independently [3]. Routing code that keeps only the top score would drop the ecommerce label. It needs a threshold for each label.

Cloudflare built Clef and Clef-flash to the interface of a rival's product, Typesafe AI's Jev [3][4]. In my view this is the most useful engineering choice in the release, because the typed answer is what downstream code depends on. Cloudflare states the Jev compatibility for its hosted models on Workers AI [4]. The same weights are on Hugging Face under Apache 2.0 [5].

Compatibility stops at Jev's feature set. Jev classifies text only, within a 32k window [11][12]. Clef's image input and 64k window have no Jev equivalent. I'd write decision steps against the subset both models share and mark image and long-context calls as Clef-only in the code.

Cloudflare says Clef currently leads the Jev Decision Index [6]. The evaluations in that table are ones Cloudflare shortlisted as important for decision-making, and Cloudflare scored the competing models itself [13]. The more useful result is on Typesafe's own eval suite: Cloudflare ran it and says Clef beat Jev in three of four areas [14]. A rival's test suite leaves less room to pick favourable ground. Both sets of scores come from Cloudflare's runs [13][14].

The 2.2-second timing covers a whole workflow. It includes fetching and rendering the page through Browser Run as well as classifying it [9][10]. gpt-oss-120b, Cloudflare's fastest general LLM, took 4.7 seconds in the same workflow and returned only two classifications [10]. The LLM took about 2.1 times as long, a gap of 2.5 seconds [1][2]. Cloudflare wrote that this is "a 2x savings in latency and results" [16], counting the extra labels as a second kind of saving. If fetch and render took the same time in both runs, the whole 2.5 seconds is model time, and the model-only ratio is larger than 2.1x [4]. For the figure to transfer, a team's inputs would need to be rendered web pages and its current baseline would need to be a general LLM.

What to watch

  • Whether Clef models tuned with Cloudflare's new RL product can be exported as weights or stay on Workers AI.
  • Whether Typesafe adds image input or a context window beyond 32k to Jev, widening the subset a fallback can cover.
  • Independent runs of the Jev Decision Index or Typesafe's suite that reproduce or contradict Cloudflare's scores.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories