BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Cloudflare's Clef-omni scores audio and video against a decision schema in one call
Cloudflare's Clef-omni takes audio and video alongside text and images and returns schema-bound decisions in one API call. Runtimewire reported that the same release cut hosted Clef-flash's context window from 64,000 to 24,000 tokens, so longer prompts now go to Clef or to self-hosted weights.
The Engineer · Build desk
What happened
- Clef-omni is a 30B-A3B mixture-of-experts model post-trained from Qwen3-Omni-30B-A3B-Instruct, with the foundation's speech-output components left unloaded at inference.
- Cloudflare cut hosted Clef-flash's input price from $0.09 to $0.038 per million tokens.
- Cloudflare says 0.24% of its Clef-flash requests carried more input tokens than the new hosted limit allows.
- Serving-layer work made hosted Clef up to 2.0 times faster, according to Cloudflare, and no new Clef weights were released for it.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability Voice and video triage that today chains a transcription model into a classifier can run as one request to one endpoint, provided Clef-omni's scores hold up on the team's own media.
- decision Text-only teams now have to weigh Clef's higher invoice-test score against Clef-omni's lower list price, and only their own evaluation data can settle it.
- constraint Clef-flash users with prompts past the new hosted limit must move up to Clef's higher price or take on serving the open weights themselves.
A Clef-omni request is one POST to the Workers AI endpoint for `@cf/cloudflare/clef-omni` [15]. The JSON body holds a `state` string, base64 data URIs in `images`, `audio` and `videos` arrays, and a `questions` object in which each key carries a type and plain-language instructions [15]. Audio goes in as WAV or MP3, video as MP4 or WebM [1]. The published sample reviews one installation. It asks whether the model and serial number label is visible in the photo, whether the unit runs without rattling or grinding, and whether the fan is running in the video [16]. The model returns a probability for each option the schema allows [14].
Inside, the model makes one prefill pass over the complete payload and generates no output tokens, according to Cloudflare [4]. Every valid option first gathers evidence from the input. Field vectors then cross-attend across the full context to compute confidence scores [5]. Runtimewire, citing the model card, describes a joint schema head that scores options across questions together [6]. I think this is the right design for picking from a fixed menu. There is no decode loop to pay for, and teams that route on a confidence threshold get a model trained for calibration. Training freezes the Qwen3 backbone, fits LoRA adapters, and pairs label-smoothed cross-entropy with Brier score calibration [7].
Cloudflare's post makes the pipeline case in one sentence: "Instead of setting up cascading pipelines of models that transcribe speech-to-text, or splitting audio and image channels from video, you can just call one model to make decisions across any modality." [2]
Cloudflare's latency medians are about 130 milliseconds for text, about 150 for images and a few hundred for audio, and it timed a 21-second video with sound at about 1.5 seconds [8]. The pass covers the whole payload [4]. For the video figure to transfer, your media has to resemble that one clip. A ten-minute support call is a much larger prefill. Runtimewire noted that the published material does not include independent testing of accuracy or latency on customer workloads [9].
Workers AI lists Clef-omni at $0.15 per million input tokens and Clef at $0.24 [17]. The multimodal model costs 37.5% less per input token [23]. The scoreboard comes from Cloudflare's own evaluations [11]. Clef-omni posts 94.8 macro-F1 on BANKING77 and 97.7 on CLINC150+OOS, and trails Clef on several other tests [10]. On the invoice-processing exact-action workflow it scores 60.2, against 64.7 for Clef and 61.8 for TypeSafe's Jev [11]. That puts it 4.5 points behind Clef and 1.6 behind Jev [22].
For a model line sold on unambiguous answers, the Jev price claim is ambiguous. Cloudflare's post says the cut makes Clef-flash cheaper than Jev [26]. Runtimewire attributes the claim to Clef-omni and adds that the comparison depends on the applicable Jev price and workload [27].
On Clef-flash, the hosted price fell about 58% [21] while the hosted window shrank 62.5% [24]. Cloudflare's share of oversized requests is measured across all of its Clef-flash traffic [19]. It transfers to a given team only if that team's prompt lengths match Cloudflare's mix. One customer sending long contracts could make up much of that tail. Publishing the share was the right call, and so was leaving the weights alone. Self-hosters keep a 256,000-token window [20], more than ten times the hosted limit [25].
What to watch
- Independent accuracy and latency results for Clef-omni on customer audio and video longer than a short clip.
- Whether Cloudflare states a Jev price basis and settles which of its models it claims undercuts Jev.
- Whether the hosted Clef-flash window moves again, or similar hosted limits appear on Clef or Clef-omni.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence45
- Adoption15
- Hype gap+35
- Incentives75
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Clef-omni accepts audio (WAV or MP3) and video (MP4 or WebM) alongside text and images.
ReportedSupportedSource: Cloudflare blog2 sources— create a free account to open themView cited source - [2]
"Instead of setting up cascading pipelines of models that transcribe speech-to-text, or splitting audio and image channels from video, you can just call one model to make decisions across any modality."
ReportedSupportedSource: Cloudflare blog post2 sources— create a free account to open themView cited source - [3]
Clef-omni's Hugging Face model card describes a 30B-A3B mixture-of-experts model post-trained from Qwen3-Omni-30B-A3B-Instruct, using the foundation's comprehension backbone and audio and vision encoders; speech-output components are not loaded for inference.
ReportedSupportedSource: Runtimewire, citing the model card2 sources— create a free account to open themView cited source - [4]
Clef-omni executes a quick prefill pass across the complete payload, scoring all modalities and valid options simultaneously; it skips output token generation, removing the overhead of transcribing or captioning incoming files.
ReportedSupportedSource: Cloudflare blog2 sources— create a free account to open themView cited source - [5]
Clef-omni uses two-stage attention routing: every valid option gathers key evidence from the input before field vectors cross-attend across the full context to compute confidence scores.
ReportedSupportedSource: Cloudflare blog2 sources— create a free account to open themView cited source - [6]
The model card describes a joint schema head that scores the options across questions together.
ReportedSupportedSource: Runtimewire, citing the model card2 sources— create a free account to open themView cited source - [7]
Clef-omni is trained by freezing the Qwen3 backbone, training low-rank adapters (LoRA), and combining label-smoothed cross-entropy loss with Brier score calibration.
ReportedSupportedSource: Cloudflare blog2 sources— create a free account to open themView cited source - [8]
Cloudflare reports median response times of about 130 milliseconds for text, about 150 milliseconds for images and a few hundred milliseconds for audio; a 21-second video with sound takes about 1.5 seconds to score.
ReportedSupportedSource: Cloudflare, as reported by Runtimewire2 sources— create a free account to open themView cited source - [9]
The latency figures are company-reported; the published material does not provide independent testing of accuracy or latency on customer workloads.
- [10]
Clef-omni posts 94.8 macro-F1 on BANKING77 and 97.7 on CLINC150+OOS, while scoring below the existing Clef model on several other tests.
ReportedSupportedSource: Cloudflare benchmark table, as reported by Runtimewire2 sources— create a free account to open themView cited source - [11]
On Cloudflare's TypeSafe workflow tests, Clef-omni scores 60.2 for invoice-processing exact actions, compared with 64.7 for Clef and 61.8 for Jev; the comparisons are Cloudflare's own evaluations.
- [12]
Cloudflare cut Clef-flash's hosted input price from $0.09 to $0.038 per million tokens.
- [13]
Cloudflare says hosted Clef is up to 2.0 times faster after serving-layer optimizations; it released no new Clef weights for that change.
- [14]
Clef returns probabilities for options supplied in a schema; the format can route a support request, classify a security event or select an action for an agent.
- [15]
A Clef-omni request is a POST to the Workers AI ai/run/@cf/cloudflare/clef-omni endpoint with a JSON body containing a state string, base64 data URIs in images, audio and videos arrays, and a questions object where each question has a type and instructions.
- [16]
Cloudflare's example asks whether the model and serial number label is visible in the photo, whether the unit sounds like it is running smoothly without rattling or grinding, and whether the fan is running in the video.
- [17]
The Workers AI pricing page lists Clef at $0.24 per million input tokens and Clef-omni at $0.15.
- [18]
Cloudflare reduced the hosted Clef-flash context window from 64,000 tokens to 24,000.
- [19]
Cloudflare says 0.24% of its Clef-flash requests exceeded 24,000 input tokens.
- [20]
Clef-flash's published model weights remain unchanged, and Cloudflare says self-hosters can use a 256,000-token context window.
- [21]
Clef-flash's hosted input price cut is about 58%.
- [22]
On the invoice-processing exact-action test, Clef-omni trails Clef by 4.5 points and Jev by 1.6 points.
- [23]
Clef-omni's listed input price is 37.5% lower than Clef's.
- [24]
The hosted Clef-flash context window is 62.5% smaller.
- [25]
The self-hosted 256,000-token window is more than ten times the 24,000-token hosted limit.
- [26]
Cloudflare says it cut the price of Clef-flash so it is now cheaper than Jev.
ReportedContestedSource: Cloudflare blog2 sources— create a free account to open themView cited source - [27]
Runtimewire reports Cloudflare says Clef-omni is now cheaper than Jev, and notes the comparison depends on the applicable Jev price and workload.
Sources
1 independent publisher whose own reporting we read for this story.
- blog.cloudflare.comIntroducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
1 article · October 9, 2026
- runtimewire.comCloudflare puts audio and video into its decision model with Clef-omni
1 article · October 9, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.