Cloudflare open-sources Clef decision models

In this article
Cloudflare open-sourced Clef and Clef-flash, two decision models hosted on Workers AI, and launched a reinforcement learning platform for fine-tuning them on customer data. What the combination changes for teams that route classification tasks to large general-purpose LLMs: cost, latency and data control.
Cloudflare released Clef and Clef-flash, open-source decision models hosted on Workers AI, plus a reinforcement learning platform for fine-tuning them on customer data. Decision models return typed outputs with probabilities instead of generated text, targeting the classification and routing steps teams currently run on large general-purpose LLMs. Clef posted a 209.3 ms median latency across 43 benchmarks, is Jev-API compatible, and its weights are on Hugging Face under Apache 2.0.
Ticket triage, intent detection, content moderation, tool selection: a large share of what teams point at large general-purpose LLMs today is not generation at all. It is classification. A model reads a state and returns a decision that code acts on. Running those steps on a general-purpose model works, but it inherits token-by-token generation, variable latency, per-output pricing and text that must be parsed and validated before anything can act on it. Cloudflare's latest release addresses exactly this layer.
What Cloudflare announced
On 1 October 2026 Cloudflare introduced Clef and Clef-flash, two decision models trained in-house and hosted on Workers AI, the company's inference platform. Both are open-sourced on Hugging Face under an Apache 2.0 license, so they can also run outside Cloudflare's infrastructure. Alongside the models, Cloudflare debuted a reinforcement learning product that lets customers fine-tune Clef on their own data. According to the announcement, Clef currently leads the Jev Decision Index, the benchmark suite defined for this model category.

A decision model does not write prose. You pass it a state, as text, JSON, images or video, and a schema of typed questions; it returns a probability for every allowed option of every question. In the support-ticket example from the announcement, a customer message is evaluated for urgency, for the team that should handle it and for a severity score, and the calling code routes the ticket, escalates or defers to a human based on those numbers. The Workers AI model page documents three question types (noul, choice and score), between 1 and 64 questions per request, and up to four embedded images per call.
Why not keep using a general-purpose LLM?
The category was defined by TypeSafe AI, which released Jev on 15 September 2026 as its first System One Model. In TypeSafe's announcement, founder Diogo Almeida argues that models optimized for human preference fit automation poorly: outputs are strings that must be parsed and validated, confidence estimates are inconsistent, and sequential sampling caps throughput. Jev instead returns type-safe structured values with calibrated probabilities, generated in parallel, with input priced at $0.042 per million tokens and output tokens free. Clef is fully Jev-API compatible and produces strictly typed outputs, so existing integrations can swap models with minimal code changes.
Latency is a major differentiator. Across the 43 evaluation benchmarks Cloudflare ran, Clef posted a median response time of 209.3 ms and Clef-flash 38.8 ms, against 524.1 ms for Jev. Laya, another competing model, is faster still at 5.8 ms median while trading away quality in the published scores. For context, TypeSafe puts end-to-end response times for frontier LLMs at 3 to 329 seconds. Because Clef runs on Cloudflare's edge GPUs, the company positions it for the hot path of an agent: decide close to the user, then hand the action to an LLM on Workers AI when generation is needed.
The announcement includes one concrete internal measurement. Cloudflare's Threat Intelligence team uses Clef, together with browser rendering, to classify website domains. Fetching, rendering and classifying a site took 2.2 seconds; the same workflow with gpt-oss-120b, the fastest general LLM in their comparison, took 4.7 seconds and returned only two classifications. For pipelines that classify many items in sequence, that gap compounds.
Benchmarks: strong in places, not uniform
On the Jev Decision Index selections published in the post, Clef scores 98.47 on BFCL (case exact) against Jev's 95.75, and 94.20 macro-F1 on BANKING77 against 79.74. Clef-flash leads BFCL at 98.76 and reaches 97.73 on the Home appliances task. The picture is not one-sided. Jev scores higher on When2Call (80.97 versus 72.37), on BRIGHT retrieval (47.52 versus 45.91) and, in TypeSafe's own four-task eval suite, on agent trace observability (71.6 versus 68.5). Clef came ahead in the other three tasks of that suite: invoice processing, customer service and security incidents.

The practical reading is that published benchmarks disagree at the margins and the ranking depends on task shape. These numbers work as a shortlist signal, not as a substitute for evaluating candidate models on your own traffic before rerouting a production classification path.
How Clef is built
Clef is a 27B multimodal model with a 65,536-token context window and a vision encoder, which distinguishes it from Jev's text-only input. Cloudflare post-trained a Qwen base model for decision tasks. The approach builds on Cloudflare's earlier experiments, which adapted DiffusionGemma to output deterministic probabilities by exposing logprobs, and on independent contributions by Matt Mastracci that strengthened DiffusionGemma support in the vLLM inference engine. At inference time the backbone runs a prefill-only pass and then scores all valid schema choices in parallel. The decision step is non-autoregressive: no intermediate text is generated token by token, which is what makes Clef significantly faster than autoregressive LLMs.
The RL fine-tuning loop and data control
The second half of the announcement is the reinforcement learning platform. Customers can fine-tune Clef for their own workloads; the service starts as a hands-on engagement with Cloudflare's forward-deployed engineer team and will later open as a self-serve platform to train and redeploy the model on Cloudflare. Cloudflare's Clef page frames the combined pitch as agents that classify, route and act without a human in the loop, tuned on the customer's own data.
Data control has two layers. On the hosted side, Cloudflare states that it does not read, store or train on requests or responses, with one explicit exception: the fine-tuning product, where using your data is the point. On the self-hosted side, the Apache 2.0 weights make local deployment possible for teams whose compliance posture rules out third-party inference. On Workers AI, the model documentation lists pricing at $0.24 per million input tokens for Clef.

Trying it: the API shape
Requests pair a state with a question schema, and answers come back keyed by question id with per-option probabilities. Long state text is truncated to fit the context window. The core of the documented Workers AI binding for JavaScript looks like this:
const response = await env.AI.run("@cf/cloudflare/clef", {
model: "clef",
state: "Checkout has been failing for every customer for the last hour.",
questions: {
urgent: { type: "noul", instructions: "Is this support request urgent?" },
team: {
type: "choice",
instructions: "Which team should handle this request?",
criteria: {
billing: "Payments, invoices, and refunds",
technical: "Outages, errors, and configuration",
sales: "Plans and upgrades",
},
},
severity: {
type: "score",
instructions: "How severe is the customer impact?",
criteria: ["No impact", "Minor", "Major", "Critical"],
},
},
});
// response.answers.urgent -> probability the request is urgent
// response.answers.team -> chosen team with per-option probabilities
// response.answers.severity -> probability-weighted scoreWhat this changes, and what to weigh
Cost: classification moves from generated-output pricing to a fixed input-token price. TypeSafe claims general LLM input pricing spans $0.20 to $10 per million tokens, with output roughly five times higher, so the gap on high-volume classification is structural.
Latency: sub-250 ms medians put decisions inside interactive loops where multi-second LLM calls force batching or queues.
Determinism: typed outputs with probabilities remove the parse-and-validate layer and shrink the set of failure modes application code must handle.
Data control: open weights cover self-hosting, the hosted service carries a no-training guarantee, and fine-tuning is a deliberate opt-in.
Inventory the classification and routing calls currently served by a general-purpose LLM.
Pilot Clef through the Jev-compatible API on a copy of that traffic and measure agreement with current outputs.
Route latency-critical decisions to Clef-flash and precision-critical ones to the larger Clef model.
Keep the LLM for generation steps, and use probability thresholds to defer low-confidence cases to humans.
The trade-off is scope. A decision model gives up open-ended generation: it classifies, scores and routes, and it needs an LLM alongside it whenever the next step is writing text or calling tools with free-form arguments. Adopting one also adds bounded engineering work: question schemas to design, probability thresholds to tune and a fine-tuning loop to govern. That is still a smaller surface than coaxing a general model into consistent behavior with prompts alone.
Decision models do not replace LLMs; they take over the decision layer, where volume is high and creativity is low. With Clef open-sourced and a hosted fine-tuning loop attached, that layer now has an option that is both self-hostable and tunable on private data. Whether it earns a place in a given stack comes down to the three numbers that motivated it: cost per decision, latency per decision, and where the data is allowed to go.
Key takeaways
- Cloudflare released Clef and Clef-flash, open-source decision models hosted on Workers AI and published on Hugging Face under an Apache 2.0 license.
- Decision models return typed outputs with probabilities instead of generated text, aimed at classification and routing inside agentic workflows.
- Across 43 evaluation benchmarks, Clef posted a 209.3 ms median latency and Clef-flash 38.8 ms, against 524.1 ms for TypeSafe's Jev.
- Benchmark results are mixed: Clef leads on BFCL and BANKING77, while Jev scores higher on When2Call, BRIGHT and agent trace observability.
- Clef is a 27B multimodal model with a 65,536-token context window, priced at $0.24 per million input tokens on Workers AI.
- The new RL platform lets customers fine-tune Clef on their own data, first with Cloudflare's forward-deployed engineers, then as a self-serve platform.
- Cloudflare states it does not read, store or train on requests or responses unless a customer opts into the fine-tuning product.
Frequently asked questions
What is a decision model?
A model that takes a state and a schema of typed questions and returns a probability for every allowed option, so application code can route, escalate or defer to a human without parsing generated text.
How much does Clef cost on Workers AI?
The Workers AI model documentation lists $0.24 per million input tokens for Clef.
Can Clef run outside Cloudflare?
Yes. The weights are published on Hugging Face under an Apache 2.0 license, so teams can run and experiment with the models on their own infrastructure.
How does Clef differ from Jev?
Both expose a compatible API. Clef adds a vision encoder for image input, a 64k-token context window versus Jev's 32k, and runs on Workers AI; Jev is TypeSafe AI's first System One Model.
What does the RL fine-tuning platform offer?
It lets customers fine-tune Clef for their own use cases. It starts as a hands-on service with Cloudflare's forward-deployed engineer team and will later become a self-serve platform to train and redeploy the model.
Does Cloudflare train on customer traffic?
Cloudflare states it does not read, store or train on requests or responses, unless you use the fine-tuning product, where your data is used to adapt the model.
When is a general-purpose LLM still the right tool?
When the task requires generating text or free-form tool calls. A decision model covers bounded classification and scoring; Cloudflare suggests combining Clef for decisions with an LLM on Workers AI for actions.












