Build2 publishersReports disagree3 min readPublished
Cloudflare's open-weight Clef models score every allowed option in a single pass
Cloudflare released Clef, open-weight 9B and 27B models that return a probability for each allowed option in a single forward pass. An agent's routing or escalation call becomes a score it can threshold, and so far the only latency figures come from Cloudflare.
The Engineer · Build desk

What happened
- Cloudflare says Clef has a vision encoder that the text-only Jev lacks, plus a 64k context window against Jev's 32k.
- Both models can be called through Workers AI or downloaded as weights from Hugging Face.
- A fine-tuning service will adapt Clef to customer data, starting with help from Cloudflare engineers, and self-service is planned with no firm date.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams already calling Jev can trial Clef without rewriting their integration, so the choice between them comes down to image input, context length and where inference runs.
- constraint The sub-40 ms case depends on Cloudflare's edge GPUs, so teams self-hosting the open weights have to measure latency again on their own hardware before relying on it.
- exposure An agent that routes or escalates on a confidently wrong score never hands the case to a human, so calibration on shifted data matters more to buyers than a benchmark accuracy figure.
- cost Customers with their own label sets will be waiting on Cloudflare engineers' availability for custom models until the self-service platform ships.
Clef takes a state and a schema of typed questions as input. The state can be text, JSON, images or video [2]. For each question, the model returns a probability for every allowed option in one forward pass [2]. It writes no free-form text, so nothing needs parsing [2]. The agent reads the scores and acts on them: it routes the support request, escalates it, or defers to a human [3].
For a support queue with a fixed set of destinations, I think this is the right design. The outcomes are a closed set, known before the call, and the schema writes them down [2]. With free-text generation, the agent has to parse an answer before it can act on it [2]. Clef returns scores against the schema, so the parse step goes away [2].
Cloudflare did not invent the format. Clef's API is compatible with Typesafe AI's Jev System One [4]. Matching a competitor's API is a polite way to concede who wrote the format. On Hacker News, Jacek Złydach wrote that "it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up" [11]. Cloudflare's engineers name two differences. Clef has a vision encoder, while Jev "only does text classification today," and Clef's context window is 64k tokens against Jev's 32k [10]. That gives it twice the room for input state [14].
Speed is the argument for the smaller model. Clef-Flash, at 9B parameters, is built for latency-sensitive decisions and posted a median of 38.8 ms in Cloudflare's benchmarks, against 209.3 ms for the 27B Clef [5]. At the median, that is about 5.4 times faster, or 170.5 ms saved per call [13]. Michelle Chen, Alex Reneau and Kevin Flansburg of Cloudflare wrote that because the models run on the company's edge GPUs, "you could put Clef into the hot path for agents to make decisions and combine that with one of our LLMs on Workers AI to take action" [15].
The 38.8 ms figure transfers only if three things hold. The call has to run on Cloudflare's hosting. Your state and schema have to be about the size of the benchmark's. And the median has to be the statistic your agent cares about. InfoQ's account reports medians only and does not describe the benchmark inputs [5]. Teams that download the weights from Hugging Face [9] run on their own hardware, and the edge-latency argument stays with Cloudflare [15].
The harder question is whether the probabilities mean anything. A per-option score lets an agent set a threshold below which it defers to a human [3]. That works only if the score drops when the model is wrong. On Reddit, bugra_sa wrote: "I'd care more about whether Clef knows when to punt than its raw accuracy score." The same user proposed testing on cases "where a false positive is much more expensive than a miss," then changing the data enough "to see when its confidence falls apart" [7]. On Hacker News, SebastianSosa wrote: "Public benchmarks are easy to cheat" [6].
Customers adapt Clef to their own data through a fine-tuning service that starts with help from Cloudflare engineers. A self-service platform is planned, and no firm date has been announced [8]. Cloudflare's own first targets are support triage and bot classification, tuned on its historical labelled data [12].
What to watch
- A release date for the self-service fine-tuning platform, which decides whether customers can tune Clef without Cloudflare engineers involved.
- Independent latency numbers for Clef-Flash that include tail percentiles and self-hosted runs, not only Cloudflare's medians.
- Published calibration tests under shifted data, of the kind bugra_sa proposed, showing whether Clef's confidence drops when it is wrong.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence50
- Adoption15
- Hype gap+35
- Incentives70
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
During Birthday Week, Cloudflare announced Clef, a set of open-weight AI models designed to choose between predefined options rather than generate text, releasing 9B- and 27B-parameter models and a platform for adapting them to specific decision-making tasks.
- [2]
Clef is a 27B multimodal model that takes a state and a schema of typed questions as input; it can process text, JSON, images or video, and returns a probability for each allowed option for every question in a single forward pass, without generating free-form text or requiring output parsing.
ReportedSupportedSource: InfoQ, citing Cloudflare blog2 sources— create a free account to open themView cited source - [3]
A decision model classifies inputs and returns typed outcomes with probabilities, which AI agents can use to decide how to act, such as routing a support request, escalating it, or deferring the decision to a human.
ReportedSupportedSource: InfoQ, citing Cloudflare blog2 sources— create a free account to open themView cited source - [4]
Cloudflare's Clef API is compatible with Typesafe AI's Jev System One model.
ReportedSupportedSource: InfoQ, citing Cloudflare blog2 sources— create a free account to open themView cited source - [5]
Clef-Flash, a 9B multimodal model designed for latency-sensitive decisions, has a median latency of 38.8 ms in Cloudflare's benchmarks, compared with 209.3 ms for the 27B Clef model.
ReportedSupportedSource: InfoQ, reporting Cloudflare's benchmarks2 sources— create a free account to open themView cited source - [6]
Public benchmarks are easy to cheat
ReportedSupportedSource: Hacker News user SebastianSosa, as quoted by InfoQ2 sources— create a free account to open themView cited source - [7]
I'd care more about whether Clef knows when to punt than its raw accuracy score. Test it on cases where a false positive is much more expensive than a miss, then change the data enough to see when its confidence falls apart. If it stays confident through that, the benchmark number doesn't mean much.
ReportedSupportedSource: Reddit user bugra_sa, as quoted by InfoQ2 sources— create a free account to open themView cited source - [8]
Cloudflare announced a fine-tuning service that lets customers adapt Clef to their own workloads using their data, initially with support from Cloudflare engineers, with a self-service platform planned for a later release; no firm date has been announced.
- [9]
The Clef models are available through Workers AI and as downloadable weights on Hugging Face.
- [10]
First, it has a vision encoder so it's able to take in images and classify visual content. This is different from Jev, which only does text classification today. Secondly, our model has a 64k context window (compared to Jev's 32k), which allows users to squeeze more input state for the model to classify against.
ReportedSupportedSource: Chen, Reneau and Flansburg of Cloudflare, as quoted by InfoQView cited source - [11]
it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up
- [12]
Cloudflare plans to fine-tune Clef for specific use cases such as support triage and bot classification, using its historical labelled data to improve accuracy and speed.
- [13]
At the median, Clef-Flash is about 5.4 times faster than the 27B Clef, a saving of 170.5 ms per call.
- [14]
Clef's 64k context window is twice Jev's 32k.
- [15]
Because they are hosted on Cloudflare's infrastructure, we're able to take advantage of our GPUs at the edge, leading to low network latency and faster decisions. This means that you could put Clef into the hot path for agents to make decisions and combine that with one of our LLMs on Workers AI to take action.
ReportedContestedSource: Michelle Chen (group product manager), Alex Reneau (principal machine learning engineer) and Kevin Flansburg (senior engineering manager), Cloudflare, as quoted by InfoQ2 sources— create a free account to open themView cited source
Sources
2 independent publishers whose own reporting we read for this story.
- dev.toCloudflare Just Shipped Two Decision Models. I Raced Them Against Jev (On OpenRouter)
1 article · October 8, 2026
- infoq.comCloudflare Open Sources Decision Models for AI Agents
1 article · October 7, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.