Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Cloudflare's Clef-omni scores audio and video against a decision schema in one call

Cloudflare's Clef-omni takes audio and video alongside text and images and returns schema-bound decisions in one API call. Runtimewire reported that the same release cut hosted Clef-flash's context window from 64,000 to 24,000 tokens, so longer prompts now go to Clef or to self-hosted weights.

The Engineer · Build desk

How we use AISend a correction

What happened

  • Clef-omni is a 30B-A3B mixture-of-experts model post-trained from Qwen3-Omni-30B-A3B-Instruct, with the foundation's speech-output components left unloaded at inference.
  • Cloudflare cut hosted Clef-flash's input price from $0.09 to $0.038 per million tokens.
  • Cloudflare says 0.24% of its Clef-flash requests carried more input tokens than the new hosted limit allows.
  • Serving-layer work made hosted Clef up to 2.0 times faster, according to Cloudflare, and no new Clef weights were released for it.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability Voice and video triage that today chains a transcription model into a classifier can run as one request to one endpoint, provided Clef-omni's scores hold up on the team's own media.
  • decision Text-only teams now have to weigh Clef's higher invoice-test score against Clef-omni's lower list price, and only their own evaluation data can settle it.
  • constraint Clef-flash users with prompts past the new hosted limit must move up to Clef's higher price or take on serving the open weights themselves.

A Clef-omni request is one POST to the Workers AI endpoint for `@cf/cloudflare/clef-omni` [15]. The JSON body holds a `state` string, base64 data URIs in `images`, `audio` and `videos` arrays, and a `questions` object in which each key carries a type and plain-language instructions [15]. Audio goes in as WAV or MP3, video as MP4 or WebM [1]. The published sample reviews one installation. It asks whether the model and serial number label is visible in the photo, whether the unit runs without rattling or grinding, and whether the fan is running in the video [16]. The model returns a probability for each option the schema allows [14].

Inside, the model makes one prefill pass over the complete payload and generates no output tokens, according to Cloudflare [4]. Every valid option first gathers evidence from the input. Field vectors then cross-attend across the full context to compute confidence scores [5]. Runtimewire, citing the model card, describes a joint schema head that scores options across questions together [6]. I think this is the right design for picking from a fixed menu. There is no decode loop to pay for, and teams that route on a confidence threshold get a model trained for calibration. Training freezes the Qwen3 backbone, fits LoRA adapters, and pairs label-smoothed cross-entropy with Brier score calibration [7].

Cloudflare's post makes the pipeline case in one sentence: "Instead of setting up cascading pipelines of models that transcribe speech-to-text, or splitting audio and image channels from video, you can just call one model to make decisions across any modality." [2]

Cloudflare's latency medians are about 130 milliseconds for text, about 150 for images and a few hundred for audio, and it timed a 21-second video with sound at about 1.5 seconds [8]. The pass covers the whole payload [4]. For the video figure to transfer, your media has to resemble that one clip. A ten-minute support call is a much larger prefill. Runtimewire noted that the published material does not include independent testing of accuracy or latency on customer workloads [9].

Workers AI lists Clef-omni at $0.15 per million input tokens and Clef at $0.24 [17]. The multimodal model costs 37.5% less per input token [23]. The scoreboard comes from Cloudflare's own evaluations [11]. Clef-omni posts 94.8 macro-F1 on BANKING77 and 97.7 on CLINC150+OOS, and trails Clef on several other tests [10]. On the invoice-processing exact-action workflow it scores 60.2, against 64.7 for Clef and 61.8 for TypeSafe's Jev [11]. That puts it 4.5 points behind Clef and 1.6 behind Jev [22].

For a model line sold on unambiguous answers, the Jev price claim is ambiguous. Cloudflare's post says the cut makes Clef-flash cheaper than Jev [26]. Runtimewire attributes the claim to Clef-omni and adds that the comparison depends on the applicable Jev price and workload [27].

On Clef-flash, the hosted price fell about 58% [21] while the hosted window shrank 62.5% [24]. Cloudflare's share of oversized requests is measured across all of its Clef-flash traffic [19]. It transfers to a given team only if that team's prompt lengths match Cloudflare's mix. One customer sending long contracts could make up much of that tail. Publishing the share was the right call, and so was leaving the weights alone. Self-hosters keep a 256,000-token window [20], more than ten times the hosted limit [25].

What to watch

  • Independent accuracy and latency results for Clef-omni on customer audio and video longer than a short clip.
  • Whether Cloudflare states a Jev price basis and settles which of its models it claims undercuts Jev.
  • Whether the hosted Clef-flash window moves again, or similar hosted limits appear on Clef or Clef-omni.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence45
Adoption15
Hype gap+35
Incentives75
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Clef-omni accepts audio (WAV or MP3) and video (MP4 or WebM) alongside text and images.

  2. [2]

    "Instead of setting up cascading pipelines of models that transcribe speech-to-text, or splitting audio and image channels from video, you can just call one model to make decisions across any modality."

    ReportedSupportedSource: Cloudflare blog post2 sources— create a free account to open themView cited source
  3. [3]

    Clef-omni's Hugging Face model card describes a 30B-A3B mixture-of-experts model post-trained from Qwen3-Omni-30B-A3B-Instruct, using the foundation's comprehension backbone and audio and vision encoders; speech-output components are not loaded for inference.

    ReportedSupportedSource: Runtimewire, citing the model card2 sources— create a free account to open themView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. blog.cloudflare.com

    1 article · October 9, 2026

    Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
  2. runtimewire.com

    1 article · October 9, 2026

    Cloudflare puts audio and video into its decision model with Clef-omni

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories