Skip to content

Build2 publishersReports disagree3 min readPublished

Cloudflare's open-weight Clef models score every allowed option in a single pass

Cloudflare released Clef, open-weight 9B and 27B models that return a probability for each allowed option in a single forward pass. An agent's routing or escalation call becomes a score it can threshold, and so far the only latency figures come from Cloudflare.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Cloudflare's open-weight Clef models score every allowed option in a single pass
Generated illustration

What happened

  • Cloudflare says Clef has a vision encoder that the text-only Jev lacks, plus a 64k context window against Jev's 32k.
  • Both models can be called through Workers AI or downloaded as weights from Hugging Face.
  • A fine-tuning service will adapt Clef to customer data, starting with help from Cloudflare engineers, and self-service is planned with no firm date.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Teams already calling Jev can trial Clef without rewriting their integration, so the choice between them comes down to image input, context length and where inference runs.
  • constraint The sub-40 ms case depends on Cloudflare's edge GPUs, so teams self-hosting the open weights have to measure latency again on their own hardware before relying on it.
  • exposure An agent that routes or escalates on a confidently wrong score never hands the case to a human, so calibration on shifted data matters more to buyers than a benchmark accuracy figure.
  • cost Customers with their own label sets will be waiting on Cloudflare engineers' availability for custom models until the self-service platform ships.

Clef takes a state and a schema of typed questions as input. The state can be text, JSON, images or video [2]. For each question, the model returns a probability for every allowed option in one forward pass [2]. It writes no free-form text, so nothing needs parsing [2]. The agent reads the scores and acts on them: it routes the support request, escalates it, or defers to a human [3].

For a support queue with a fixed set of destinations, I think this is the right design. The outcomes are a closed set, known before the call, and the schema writes them down [2]. With free-text generation, the agent has to parse an answer before it can act on it [2]. Clef returns scores against the schema, so the parse step goes away [2].

Cloudflare did not invent the format. Clef's API is compatible with Typesafe AI's Jev System One [4]. Matching a competitor's API is a polite way to concede who wrote the format. On Hacker News, Jacek Złydach wrote that "it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up" [11]. Cloudflare's engineers name two differences. Clef has a vision encoder, while Jev "only does text classification today," and Clef's context window is 64k tokens against Jev's 32k [10]. That gives it twice the room for input state [14].

Speed is the argument for the smaller model. Clef-Flash, at 9B parameters, is built for latency-sensitive decisions and posted a median of 38.8 ms in Cloudflare's benchmarks, against 209.3 ms for the 27B Clef [5]. At the median, that is about 5.4 times faster, or 170.5 ms saved per call [13]. Michelle Chen, Alex Reneau and Kevin Flansburg of Cloudflare wrote that because the models run on the company's edge GPUs, "you could put Clef into the hot path for agents to make decisions and combine that with one of our LLMs on Workers AI to take action" [15].

The 38.8 ms figure transfers only if three things hold. The call has to run on Cloudflare's hosting. Your state and schema have to be about the size of the benchmark's. And the median has to be the statistic your agent cares about. InfoQ's account reports medians only and does not describe the benchmark inputs [5]. Teams that download the weights from Hugging Face [9] run on their own hardware, and the edge-latency argument stays with Cloudflare [15].

The harder question is whether the probabilities mean anything. A per-option score lets an agent set a threshold below which it defers to a human [3]. That works only if the score drops when the model is wrong. On Reddit, bugra_sa wrote: "I'd care more about whether Clef knows when to punt than its raw accuracy score." The same user proposed testing on cases "where a false positive is much more expensive than a miss," then changing the data enough "to see when its confidence falls apart" [7]. On Hacker News, SebastianSosa wrote: "Public benchmarks are easy to cheat" [6].

Customers adapt Clef to their own data through a fine-tuning service that starts with help from Cloudflare engineers. A self-service platform is planned, and no firm date has been announced [8]. Cloudflare's own first targets are support triage and bot classification, tuned on its historical labelled data [12].

What to watch

  • A release date for the self-service fine-tuning platform, which decides whether customers can tune Clef without Cloudflare engineers involved.
  • Independent latency numbers for Clef-Flash that include tail percentiles and self-hosted runs, not only Cloudflare's medians.
  • Published calibration tests under shifted data, of the kind bugra_sa proposed, showing whether Clef's confidence drops when it is wrong.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence50
Adoption15
Hype gap+35
Incentives70
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    During Birthday Week, Cloudflare announced Clef, a set of open-weight AI models designed to choose between predefined options rather than generate text, releasing 9B- and 27B-parameter models and a platform for adapting them to specific decision-making tasks.

  2. [2]

    Clef is a 27B multimodal model that takes a state and a schema of typed questions as input; it can process text, JSON, images or video, and returns a probability for each allowed option for every question in a single forward pass, without generating free-form text or requiring output parsing.

    ReportedSupportedSource: InfoQ, citing Cloudflare blog2 sources— create a free account to open themView cited source
  3. [3]

    A decision model classifies inputs and returns typed outcomes with probabilities, which AI agents can use to decide how to act, such as routing a support request, escalating it, or deferring the decision to a human.

    ReportedSupportedSource: InfoQ, citing Cloudflare blog2 sources— create a free account to open themView cited source

Sources

2 independent publishers whose own reporting we read for this story.

  1. dev.to

    1 article · October 8, 2026

    Cloudflare Just Shipped Two Decision Models. I Raced Them Against Jev (On OpenRouter)
  2. infoq.com

    1 article · October 7, 2026

    Cloudflare Open Sources Decision Models for AI Agents

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories