Cloudflare has put a new open model on Hugging Face: Clef-Omni, a 30B-A3B mixture-of-experts that reads text, images, audio and video and answers structured questions with probabilities instead of prose. It is the third member of Cloudflare's Clef family of decision models, and it is the one built on an omni backbone with an audio encoder, so it can listen to an audio clip as well as read images and video before it decides. The model card lists an Apache-2.0 license and a Jev-compatible API, which means it can drop into existing decision-model pipelines.
If you are new to the category, start with our guide to what decision models are. Short version: instead of generating an answer and parsing it, you ask typed questions and read back a probability for each option.
Quick answers
| Question | Answer |
|---|---|
| What is it? | A multimodal decision model: state plus typed questions in, per-option probabilities out. |
| Size and architecture | 30B-A3B MoE, listed as 35B parameters in bfloat16. |
| Base model | Qwen3-Omni-30B-A3B-Instruct, post-trained by Cloudflare. |
| License | Apache-2.0, following the base model. |
| Inputs | Text, JSON, images, audio, video (up to 64,000 tokens by default). |
| GPU needs | About 64 GB in bfloat16; tested on one H200. |
| API | Jev and SystemOne compatible. |
| Serving | SGLang support is listed as coming soon. |
Where Clef-Omni fits: the Clef family
Cloudflare announced the first Clef models on October 1, 2026, about two weeks after TypeSafe launched Jev, according to The Register. That coverage describes Clef as post-trained on Qwen3.8-27B and the smaller Clef-flash on Qwen3.5-9B, both able to answer yes/no, multiple-choice and ranking questions. The Register reported pricing of $0.24 per million tokens for Clef on Workers AI, against $0.042 for Jev, and noted that Cloudflare's scores were self-reported. Clef-Omni follows as the audio-and-video variant.
If Jev is new to you, our write-ups of the Jev and System One launch and the Jev speed and cost fact-check give the baseline that Clef is measured against. Cloudflare is also not alone: OpenAI has its Decisions API, Perplexity shipped pplx-decider, and Strands released an open 2B decider. Our ranked comparison lives in the top 10 decision models.
A signpost with three arrows, one highlighted, picturing how Cloudflare Clef-Omni picks one option per typed question
How Clef-Omni works
The model card describes four moving parts.
- Record. You pass a JSON record with a
state(any JSON value or string), optionalimages,audioandvideos, and aquestionsmap. - Backbone. The Qwen3-Omni thinker reads the state, including media, and produces final hidden states. Video files are sampled at 2 frames per second and, when every video has a soundtrack, the audio is heard alongside the frames.
- Joint schema head. A small transformer head routes evidence from the state to each question and scores all options of all questions together.
- Output. One logit per allowed option per question. A softmax per question yields probabilities.
Each question is one of three types: noul (true or false), choice (named options with descriptions) and score (ordered options, returning an expected score). Because there is no generation step, there is nothing to parse and the output shape is guaranteed. That is the whole pitch of this category, and the reason latency stays predictable: one forward pass answers every question. Our explainer on System One models walks through why this is different from a chat model.
Try it: the card's example
This is the usage snippet from the model card, shortened. It checks an invoice and asks two questions.
import sys, torch
from huggingface_hub import snapshot_download
path = snapshot_download("Cloudflare/clef-omni")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model
model, processor = load_release_model(path, device="cuda")
record = {
"state": {"invoice": {"vendor": "Acme", "total": 1250.0, "status": "overdue"}},
"questions": {
"status": {"type": "choice", "instructions": "What is the invoice status?",
"criteria": {"paid": "Paid.", "overdue": "Past due.", "draft": "Not sent."}},
"large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
},
}
Then call encode_record, collate_records and run the model under torch.inference_mode(). The card tested torch 2.11 and transformers 5.10.2. For audio or video, add audio and videos lists to the record. The card's example asks whether glass is breaking in a call recording and whether a dashcam clip shows a collision.
A Jev-style call is also supported through the systemone helper, which takes a /v1/systemone request body and returns model, answers and usage. If you already run Jev in production, the card's claim of full API compatibility suggests swapping the model name is the migration path; test it first. For local runtimes, see our note on llama.cpp support for decision models.
Benchmarks: where it wins and where it does not
The numbers below come from Cloudflare's internal run of Decision Index 0.2.1 as published on the model card. Treat them as vendor-reported.
| Benchmark | Clef-Omni | Clef | Clef-flash | Jev |
|---|---|---|---|---|
| MMLU (accuracy) | 92.7 | 90.3 | 91.8 | 91.7 |
| CRUXEval (accuracy) | 88.8 | 86.7 | 86.1 | 73.0 |
| CLadder (accuracy) | 97.2 | 94.0 | 97.7 | 72.6 |
| BANKING77 (macro-F1) | 94.8 | 94.2 | 90.9 | 79.7 |
| GPQA Diamond (accuracy) | 47.4 | 48.0 | 51.0 | 78.3 |
| RAGTruth (hallucination F1) | 42.0 | 79.4 | 35.6 | 76.5 |
| ForecastBench (Brier, lower is better) | 11.7 | 13.9 | 10.6 | 17.4 |
Some takeaways:
- Strong on classification-style tasks. Intent and domain classification (BANKING77, CLINC150+OOS at 97.7) and logic tasks like CLadder are strong.
- Weak on hard reasoning and hallucination detection. GPQA Diamond and RAGTruth are well behind Jev and the dense Clef. If you are building a hallucination judge, the dense Clef is the better pick from this family.
- Mixed on workflows. On Typesafe Evals, Clef-Omni scores 60.2 exact actions on invoice processing versus 61.8 for Jev, and 71.6 on customer service versus 76.0 for Jev. It is not a clear upgrade over what you may already run.
- Audio and video are the point. Most Decision Index rows are text. The card gives no audio or video benchmark, so the multimodal strength is unquantified in the published results.
Two shapes on matching plinths joined by a dotted line, a visual for comparing Cloudflare Clef-Omni with Clef, Clef-flash and Jev
Clef-Omni versus the rest: which should you pick?
| If you need | Pick | Why |
|---|---|---|
| Decisions over audio | Clef-Omni | The only Clef variant with an audio encoder (The Register reports Clef also handles images and video) |
| Fast, cheap text classification | Clef-flash | Smaller dense model |
| Best accuracy on text and image | Clef | Leads the family on several rows |
| Hosted with no GPU | Check Workers AI listings | The card documents self-hosting; hosted options are on Cloudflare's platform |
| Hallucination or faithfulness checks | Clef or Jev | Clef-Omni scores 42.0 on RAGTruth |
What people will ask next
Is it really "open source"? The weights and loading code are Apache-2.0 on Hugging Face, which is what most developers need to self-host, fine-tune and ship commercially. The training datasets are not published, so it is open weights rather than fully open data. That mirrors its base model, Qwen3-Omni, whose license the card says it follows.
Why a decision model instead of a chat model with JSON mode? A chat model generates tokens one at a time, can drift from the schema, and gives you a confidence you have to guess at. A decision model scores the allowed options directly, so the answer is always one of the options you defined and comes with a probability you can threshold. Our structured output and JSON mode guide covers the generation-side alternative and where it still wins, such as free-text explanations.
What does the joint head buy you? Scoring all questions together lets one pass answer many questions about the same state. If you ask five questions about a support call, the audio is processed once. That is a large saving over five separate calls, and it keeps the answers consistent because they share the same evidence.
How should I set thresholds? Use the probabilities as routing signals. For example, auto-approve above 0.95, send 0.6 to 0.95 to a cheaper review step, and escalate below 0.6 to a human. Calibrate these numbers on a labeled sample of your own traffic, because published benchmark accuracy does not tell you how well-calibrated the probabilities are on your data.
What about latency? Cloudflare has not published latency for Clef-Omni on the card. Expect it to be slower than the 9B-class Clef-flash because a 30B-A3B model activates about 3B parameters per token but still has to load all 35B parameters into memory, and audio and video inputs add many tokens. Measure on your hardware before you commit.
A practical evaluation plan
If you want to decide whether Clef-Omni deserves a slot in your stack, a one-day test is enough.
- Collect 200 labeled examples from real traffic, including a spread of easy and borderline cases and at least 40 with audio or video.
- Define your questions as typed schemas. Start with
noulquestions because they are the easiest to score and calibrate. - Run Clef-Omni, Clef-flash and your current model on the same set and compare exact-match accuracy, not just average confidence.
- Check calibration. Bucket answers by predicted probability and compare to the actual hit rate in each bucket.
- Measure cost per decision. Include GPU hours for self-hosting against per-token pricing for hosted models.
- Review the misses by hand. Look for patterns such as failures on noisy audio or long videos, and decide whether a fallback is needed.
This mirrors the approach in our guide to classifier feature engineering and calibration, which explains why calibration matters more than headline accuracy when you act on probabilities.
Limits and open questions
- Self-reported scores. The Register noted Cloudflare's earlier numbers had not been reproduced on the official Decision Index. The same caution applies here.
- Hardware cost. 64 GB of GPU memory means an H200 or a pair of smaller cards, far above the 9B-class decision models.
- No generation. You cannot ask it to explain a decision. It returns probabilities only.
- Weights note. The card says the base model's speech-output weights are included unchanged but unused, so the download is larger than the decision path needs.
- Serving. SGLang support is "coming soon" per the card; for now you run the provided Python code.
- Training data. The card does not describe the post-training datasets. The Register reported a Cloudflare product manager saying the datasets for the Clef release are not public, so "open" here means weights, not data.
What this means for builders
Audio and video decisions have until now meant a pipeline: transcribe or caption, then classify. A single model that hears the call and watches the clip removes a stage and a failure mode. Good early uses are call-center triage, claims review with photos and recordings, content moderation of short clips, and safety checks on agent screen recordings. As always with this category, calibrate the confidence thresholds on your own data before you trust a probability, and keep a human in the loop for high-stakes calls. Our post on cheap verification checkpoints in agent pipelines shows where a decision model slots in. If your agents take real-world actions, AgentBeam, the agent security platform from the explainx.ai team, is built to stop agents before they take dangerous actions.
Related reading
- What are decision models?
- What is a System One model?
- Top 10 decision models ranked
- OpenAI Decisions API vs Jev
- Strands 2B open decision model
- llama.cpp decision model support
Specs and scores are accurate as of October 9, 2026 and come from the Hugging Face model card; check it for changes.
