Cloudflare

Cloudflare's Clef Bets Decisions, Not Text Generation, Belong in the Agent Hot Path

Cloudflare ships Clef, an open-source decision model with an RL fine-tuning service, cutting classification latency versus general LLMs.

Cloudflare's Clef Bets Decisions, Not Text Generation, Belong in the Agent Hot Path — article cover

Decision models got a lot of attention after Typesafe AI’s Jev landed, and on October 1, 2026, Cloudflare answered with its own family: Clef and Clef-flash, two open-source decision models on Workers AI, plus a reinforcement learning fine-tuning service built on top of them.

What a decision model actually buys you

The category is narrow on purpose. Instead of generating text or tool calls, a decision model returns typed answers with probabilities. Feed it a customer support message and ask whether it’s urgent and which team should own it; your code routes, escalates, or defers to a human based on those numbers. Cloudflare’s own example is domain classification for its Threat Intelligence team: Clef fetched, rendered, and classified a website in 2.2 seconds, while the general model gpt-oss-120b took 4.7 seconds and returned only two classifications.

That’s the whole pitch. You trade open-ended generation for deterministic, schema-bound outputs that run fast enough to sit in an agent’s hot path.

How Clef differs from Jev

Clef is API-compatible with Jev, so swapping it in should be a config change, but the model has three differences worth noting:

  • A vision encoder, so it can classify images, which Jev (text-only today) can’t.
  • A 64k context window versus Jev’s 32k.
  • Latency: across 43 benchmarks, Clef-flash posted a median of 38.8ms versus Jev’s 524.1ms. The larger Clef came in at 209.3ms.

On quality, it’s mixed rather than dominant. Cloudflare reports Clef leading the Jev Decision Index and beating Jev in 3 of 4 Typesafe eval workflows, but Jev still wins on some benchmarks like When2Call and BRIGHT retrieval. Read the table before assuming a clean win.

Both models are on Hugging Face under Apache 2.0, which matters if you want to run classification locally rather than through a hosted API. I’ve argued before that open model weights change how builders make deployment decisions, and a small classifier you can self-host is exactly the kind of workload where that choice is easy to justify.

The training approach, briefly

Clef uses Qwen as its backbone — Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash — kept frozen, with a routing head and rank-256 low-rank adapters trained on top. At inference it runs a prefill-only pass and scores valid schema choices in parallel. Because nothing is generated token by token, there’s no autoregressive latency to pay. Training combined label-smoothed cross-entropy, a Brier loss for probability calibration, and a reinforcement learning objective called RLCD that gives partial credit to adjacent choices and penalizes distribution shift.

The fine-tuning service is the part to watch

The RL product is where this gets interesting for teams with real labeled data. Cloudflare’s plan: start hands-on with a forward-deployed engineer team, then graduate to a self-serve platform where you capture data, fine-tune, and redeploy on Cloudflare. The plumbing is mostly existing pieces — AI Gateway to build datasets from your traffic, Workers AI for rollouts, Containers as an RL sandbox, a new Trainer service for weight updates, and BYO Model to redeploy.

The tradeoff is familiar: fine-tuning narrows a model to your domain, giving up some general performance for accuracy where it counts. If your agent already routes tickets or triages submissions, a domain-tuned decision model could replace both a fragile prompt and a general LLM call.

One caveat: the fine-tuning platform isn’t self-serve yet, and the source announcement doesn’t specify pricing or a timeline for the self-serve version. If you want to try the concept now, the hosted models and Hugging Face weights are available today — start with a classification step your agent already does badly, and measure whether typed probabilities beat your current prompt-and-parse approach.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

Found this useful?

Support more practical AI articles, tutorials, and build notes.

Buy us a coffee
SHAREXEMAIL