Most decision models — the ones that classify, route, and score instead of generating text — have been text-only. Cloudflare’s new Clef-omni, announced October 9, 2026, changes that: it accepts audio (wav/mp3) and video (mp4/webm) alongside text and images in a single API call. No more cascading speech-to-text pipelines or splitting video into audio and frame channels before a decision model ever sees the input.
How it works under the hood
Clef-omni is built on the Qwen3-Omni-30B-A3B-Instruct mixture-of-experts backbone, which already handled text, imagery, audio, and video. Cloudflare kept the comprehension side and dropped the text-to-speech output components. Because Clef models are not LLMs, they skip output token generation entirely: the model runs a prefill pass over the full payload, scores all modalities at once, and pulls candidate values from internal embeddings using a two-stage attention routing scheme with a built-in lexical grammar for schema-constrained scoring.
The numbers reflect that efficiency. Median latency is about 130 ms for text decisions, 150 ms for images, a few hundred milliseconds for audio, and roughly 1.5 seconds for a 21-second video clip with sound.
The benchmark picture is mixed
Against Jev (TypeSafe’s decision model), Clef-omni leads on several benchmarks — BFCL case exact at 98.2 versus 95.75, and BANKING77 macro-F1 at 94.8 versus 79.74. But it trails its own siblings and Jev elsewhere: Clef-flash hits 97.73 on Home appliances where Clef-omni scores 69.3, and Jev wins When2Call at 80.97 versus Clef-omni’s 63.3. The TypeSafe workflow evals show tighter spreads, with Clef-omni generally a few points behind its text-only variants.
The practical read: adding modalities costs some accuracy on text-heavy tasks. If your workflow is text-only, the original Clef or Clef-flash is still likely the better pick. Clef-omni earns its slot when input arrives as media.
This is a familiar pattern for anyone comparing retrieval scorecards — your retrieval benchmark may be judging yesterday’s results — benchmark fit matters as much as headline numbers.
The pricing and context-window trade-off
Clef-flash dropped from $0.09 to $0.038 per million input tokens, undercutting Jev. Clef stays at $0.24, and Clef-omni launches at $0.15.
The catch: the hosted Clef-flash context window is now 24k, down from the previously advertised 64k. Cloudflare says only 0.24% of requests exceeded 24k input tokens, so most callers won’t notice. If you do run long contexts, either switch to Clef (still 64k) or self-host — the Hugging Face weights are untouched and support 256k. It’s a reminder that hosted pricing and open weights can diverge, and your migration path is usually one config change.
Serving got faster too
Clef on Workers AI now runs 1.7× to 2.0× faster at the median, mostly from infrastructure work — notably a move to SGLang, with upstream pull requests landing in SGLang 0.5.22. No new weights shipped, so self-hosted setups benefit too by upgrading their serving stack.
What this changes for builders
Cloudflare reports internal teams already using Clef for spam issue detection in its public GitHub docs repo, phishing moderation in its CMS, PII scanning, and malicious domain detection. The interesting shift is that classification moved into the model layer: where zero-shot classifiers once needed a specialized ML team and a tuned corpus, you can now call a schema-bound decision model directly.
Clef is Jev-API compatible and available via AI Gateway, so trying it means changing the model ID. If you have media in your decision workflows — support calls, security footage, screencasts — Clef-omni is worth a real test; for pure-text routing, stick with Clef or Clef-flash and keep the 24k limit in mind.
