Cloudflare Open-Sources Clef as Decision Models Begin Competing to Become the Gateway to AI Agents

On October 1, Cloudflare released two open-source decision-making models based on Qwen, Clef and Clef-flash, both compatible with the Jev-API, and launched a reinforcement learning product for business-specific fine-tuning. They target high-frequency decision-making tasks such as classification, routing, and moderation, rather than chat or content generation.
Cloudflare Turns Models into “Smart If Statements”
On October 1 local time, Cloudflare released two open-source multimodal decision models, Clef and Clef-flash, and deployed them on Workers AI. Both are based on Qwen, targeting higher-accuracy and lower-latency decision tasks respectively; their interfaces are compatible with the Jev API, allowing developers to try them with an experience close to swapping model endpoints. Cloudflare also launched a reinforcement learning fine-tuning product that allows customers to customize models for their own business scenarios.
This is not another general-purpose LLM trying to compete for a place on the chatbot leaderboard. Clef targets the large number of repetitive, clearly bounded decisions in agent workflows: which team a ticket should be assigned to, whether a piece of content violates policy, whether a tool call should be allowed, or whether a request should be escalated to a human. It outputs structured choices, scores, or probabilities that downstream programs can consume directly, instead of first generating an explanation and then having another piece of code laboriously extract the answer from the text.

It is accurate to think of them as “smart if statements”: traditional rules can handle explicit conditions, but become brittle when faced with natural language, ambiguous intent, and context; general-purpose LLMs can process this information, but often generate much more content than the business logic actually needs. Decision models aim to occupy the middle ground by restricting model judgments to a defined set of questions and result types, leaving ordinary code to determine what happens next.
Two Sizes for Two Cost Profiles
Clef is based on Qwen3.8-27B, while Clef-flash is based on Qwen3.5-9B. The former is the accuracy-oriented version, suited to tasks where errors are costly and contextual information is extensive. The latter targets latency-sensitive “hot paths,” such as intent classification or tool selection that every message must pass through. Cloudflare lists a 64K-token context window for the Workers AI versions.
This division of labor has practical value. Model selection should not be based solely on parameter count or a single aggregate score: misrouting a customer-service ticket may result in several additional hours of waiting, while a difference of a few dozen milliseconds in an agent’s decision about whether to proceed before each tool call can accumulate across the entire chain. Handing every decision to the largest model makes costs and response times difficult to control; replacing everything with a small model may sacrifice accuracy on complex inputs. Two models of different sizes give developers room to tier decisions by risk, but whether the tradeoff is worthwhile still has to be calculated using their own traffic, hardware, and cost of errors.
Cloudflare says that, on the benchmarks it published, Clef performed competitively on several classification, tool-calling, and business-decision tasks. The release materials also provide a latency comparison with Jev: Clef has a median latency of approximately 209 milliseconds, Clef-flash approximately 39 milliseconds, and Jev approximately 524 milliseconds. Based on these figures, Clef-flash was about 13 times faster in that test.
These figures are worth noting, but they should not be treated as a speed guarantee for every production environment. They come from vendor-published tests, and the results are affected by request length, concurrency, deployment location, hardware, and the composition of the test tasks. In particular, latency figures can conceal peak-period behavior when only the median is considered. Production evaluation should at minimum also examine p95 latency, throughput, failure rates, and the business losses caused by incorrect decisions. A model being fast does not mean its decision quality is sufficient; a high average score does not mean it is reliable for every category of request.
Jev-API Compatibility Reduces Migration Costs
Jev is a type of decision-model interface developed by TypeSafe AI. Rather than aiming to generate natural language, it accepts a set of typed questions and returns results that programs can process directly. Common questions include selecting one option from a given set, assigning a score on a scale, and answering a yes-or-no question. A single call can handle multiple such judgments at once, making the interface suitable for ticket routing, content moderation, information extraction, risk scoring, and agent guardrail checks.
Clef is compatible with the Jev API, so applications already built around this type of interface can conduct relatively straightforward side-by-side tests: retain the question definitions and business logic, switch the model or service endpoint, and compare accuracy, latency, and cost. This lowers the engineering barrier to evaluating alternative models and means decision models do not first have to convince developers to rewrite an entire application before being considered.
However, “compatible” does not mean that all behavior is identical. Different models may perform differently in confidence calibration, option boundaries, and handling abnormal inputs. The same prompts and business thresholds may also need to be revalidated after switching models. This is especially important when downstream code maps probability thresholds to automatic refunds, bans, or high-risk operations: identical interface fields do not imply identical decision risks. During migration, teams should focus on boundary cases, the handling of refusals or “other” options, and whether the probability distribution is suitable for existing thresholds, rather than merely confirming that requests return successfully.
Multimodality Is More Than Adding Image Understanding to a Decision Model
Clef can also process images. Cloudflare says the model is equipped with a vision encoder and can make judgments using image inputs. Compared with decision solutions that process text only, it can incorporate screenshots, product images, photos of forms, or page states into its decisions. For agents, this means visual signals such as “What is currently displayed on the page?” and “Does the uploaded document contain the required information?” can potentially enter the same structured decision pipeline directly, without always first calling a vision model to generate a description and then having another model classify it.
The 64K context window also provides more room for business state inputs. A system can provide a longer conversation history, rule summary, user attributes, and workflow state within the same judgment. But a larger window is not always better: irrelevant history increases processing overhead and may cause key fields to get lost. Context engineering remains important. Inputs should contain state that is directly relevant to the current decision, verifiable, and appropriate to the user’s permissions, rather than simply passing the entire database to the model.
Multimodal capabilities also require separate testing. Image quality, cropping, text density, and domain differences can all change the results. Performance on text-classification benchmarks cannot be used to infer reliability in document verification, interface recognition, or product inspection. For financial, identity, or security-related decisions, teams should retain human review, audit logs, and clearly defined failure fallback paths.
Open Weights Plus Business Fine-Tuning: The Real Value Lies in the Feedback Loop
Cloudflare has released Clef’s weights under the Apache 2.0 license. Developers can obtain them from Hugging Face and experiment locally or in their own environments. For teams with requirements around data residency, request isolation, or deployment control, open weights offer more options than simply calling a hosted API. Cloudflare also says that hosted models do not read, store, or use requests and responses for training by default unless customers opt into its fine-tuning product. Before using the service, teams should still review the terms of service, data-processing agreements, and deployment configuration.
The newly launched reinforcement learning product shifts the focus from “switching to a more capable general-purpose model” to “adapting decisions to a specific business.” Customer-service routing, internal approvals, and risk reviews often have their own labeling systems, escalation rules, and costs of misclassification. Public benchmarks cannot fully substitute for these business standards. Fine-tuning on real workflows could, in theory, reduce the model’s dependence on general-corpus patterns and bring its outputs closer to an organization’s own decision boundaries.
Fine-tuning, however, does not automatically make business decisions correct. Label consistency in the training data, coverage of minority-class examples, reward-signal design, and the quality of online feedback all affect the final result. If the business rules themselves are ambiguous, the model will only reproduce that ambiguity more consistently. If the reinforcement learning reward only encourages “rapid closure,” the model may also become overly inclined toward automation. A genuinely usable process should first define which decisions can be executed automatically and which must be escalated, then validate the new model through offline replays, shadow traffic, and small-scale canary deployments while continuously monitoring the distribution of errors.
Clef Is Worth Testing, but It Is Not a General-Purpose Replacement
Clef’s product positioning is sound: agents need not only models that can reason and generate, but also a decision layer that is inexpensive, fast, and produces outputs that can be executed directly. Jev API compatibility reduces integration friction, open weights expand deployment options, and visual inputs plus business fine-tuning extend the model toward real-world workflows. For teams building multi-model agents, especially those already handling large volumes of classification, routing, and moderation calls, Clef belongs on the evaluation list.
It also has clear boundaries. Decision models are suited to making judgments among known states and options. They will not replace general-purpose models that are needed for open-ended generation, code writing, long-chain reasoning, or proactively querying external information. A more robust system architecture is to have a general-purpose model handle complex understanding and content generation, a decision model handle high-frequency, structured, and measurable branches, and business code enforce permissions, validation, and fallback logic.
The next three things worth watching are whether independent evaluations can reproduce Cloudflare’s published accuracy and speed results; whether the fine-tuning product can deliver measurable gains using enterprises’ own labels and data conditions; and whether the hosted version and self-deployed weights establish clear tradeoffs among performance, privacy, and operational cost. Clef’s release adds an open-source option to the specialized field of decision models, but production data will have to show whether it can become a standard component of agent workflows.
References
- IT Home: Cloudflare Launches Clef, an Open-Source Multimodal Decision Model Based on Qwen — Information on the release date, model foundation, and core capabilities.
- Hugging Face: Cloudflare/clef — Clef’s open-source model weights and related model information.
- Juejin: What Is Jev? TypeSafe AI’s System One Model — Explanation of Jev’s decision-interface types and typical application scenarios.



