On October 1, 2026, Cloudflare announced Clef and Clef-flash, decision models that return probabilities for predefined options instead of generating text. They are available on Workers AI, and the trained weights are released under the Apache 2.0 license.
Rather than generating an answer token by token like a typical large language model (LLM), Clef evaluates options directly: which department to route a request to, whether to continue processing, or whether to ask a human to check. The aim is to cut the time AI agents spend waiting for long answers to be generated each time they need to decide how to branch.
In Cloudflare's measurements, the median response time of the lightweight Clef-flash was 38.8 milliseconds. However, on some evaluation items there is a large accuracy gap between it and the larger Clef. To build fast decisions into real operations, you need to work out which tasks Clef-flash can handle with sufficient accuracy, and under what conditions fine-tuning on customer data is available.
Returning probabilities for each option instead of writing text
What you give Clef is a combination of the situation to be judged and the questions you want answered.
For example, for an inquiry saying "errors keep occurring during checkout," you can pose several questions at once, such as:
- Does this need urgent handling?
- Should it go to billing or to technical support?
- How large is the impact on the customer?
What comes back is not an explanatory text but probabilities or scores for the options you defined in advance. The calling program can use the result to route automatically to a department, continue processing, or send the case back for human review.
There are three question formats: noul for true/false judgments, choice for selecting among named candidates, and score for rating ordered levels.
Cloudflare says Clef is compatible with the System One API used by Typesafe AI's "Jev," making it easy to switch from existing Jev implementations.
According to Cloudflare's technical explanation, Clef is built on Qwen3.8-27B and Clef-flash on Qwen3.5-9B.
A typical LLM processes the input and then generates an answer one token at a time. Clef, by contrast, uses Qwen to process the entire input and then evaluates the permitted options in parallel with a dedicated mechanism.
In other words, instead of generating a rationale as text and extracting an answer from it, the model derives the probability of each option directly from the information obtained internally. Skipping text generation shortens the time to a decision. The models can also take images and video as input.
During training, the weights of the underlying Qwen were frozen, and the part that makes decisions, along with an adapter that tunes the model with a small number of additional parameters, was trained.
Cloudflare also says it worked on probability calibration, which brings predicted probabilities close to actual accuracy, so that an answer given with 80% confidence is right roughly 80% of the time.
The goal is not merely to fix the answer format but to make the returned probabilities easier to use for operational decisions such as whether to process automatically or hand back to a human.
However, being able to output probabilities does not in itself guarantee that a judgment is correct.
The faster Clef-flash needs to be matched to the right tasks
Cloudflare published median response times of 209.3 ms for Clef, 38.8 ms for Clef-flash, and 524.1 ms for Jev.
At the 95th percentile, which reflects the slower end, Clef-flash is also the shortest.
| Response time | Clef | Clef-flash | Jev |
|---|---|---|---|
| Median | 209.3 ms | 38.8 ms | 524.1 ms |
| 95th percentile | 238.6 ms | 122.4 ms | 536.0 ms |
Source: Cloudflare's published figures. These are Cloudflare's own evaluations and do not guarantee the same response times for users' inputs or execution environments.
For workloads where you want the shortest possible wait, Clef-flash looks attractive.
But results change substantially depending on the type of task evaluated.
In the macro-F1 on CLINC150+OOS listed in the announcement, Clef scored 97.43 while Clef-flash scored 66.77. In this evaluation, which classifies input into many categories, the larger Clef is far ahead.
On the other hand, in BFCL's per-case exact-match rate, Clef-flash scored 98.76 and Clef 98.47, so the lighter model was slightly higher.
There are also large gaps in evaluations not shown in the excerpted table in the announcement.
According to the full results on the model card, the F1 on RAGTruth, which detects errors in generated content, was 79.4 for Clef and 35.6 for Clef-flash. The difference is 43.8 points.
These are the results of Cloudflare's internal evaluation on Decision Index 0.2.1, checked on October 2, 2026. They are not numbers reproduced by an independent third party, and the public table does not show the sample sizes or statistical error for each task.
Therefore, the 43.8-point gap will not necessarily appear in the same way for other tasks or for use in Japanese.
Still, if you were to entrust Clef-flash with checking whether content generated by an AI agent contains errors, this gap is hard to ignore.
A model that is good enough for quickly routing inquiries is not necessarily good enough for finding errors in generated content.
Answering in a fixed format and how far its judgments can be trusted are separate questions.
In comparisons with competitors, too, Clef is not better on every evaluation.
On When2Call accuracy, Jev scored 80.97, Clef 72.37, and Clef-flash 65.58.
So rather than deciding to use the fastest model for every judgment, it is more useful for deployment decisions to choose evaluations close to the work you actually want to delegate and to check misses and misjudgments on your own data.
Weights are open, training data is not
According to the official Workers AI documentation, input pricing per million tokens is $0.24 for Clef and $0.09 for Clef-flash.
Both handle a context of up to 65,536 tokens. These are model usage prices as of October 2, 2026, and do not include the cost of surrounding services or fine-tuning.
Because the trained weights are public, you can also run the models in your own environment without using Workers AI.
However, just because the 27-billion-parameter Clef and the 9-billion-parameter Clef-flash take little time to make a decision does not mean they need little GPU memory to run.
In response to The Register's reporting, Michelle Chen, who leads AI Platform at Cloudflare, explained that with a 64k context and one request processed at a time, Clef-flash needs at least 41 GB of VRAM and Clef at least 85 GB.
These are values given for specific execution conditions, and do not mean the same capacity is the minimum for shorter inputs or quantized models.
Chen also acknowledged to the publication that the datasets used for training have not been released.
Cloudflare describes Clef as "open source," but being able to use trained weights under Apache 2.0 is distinct from being able to reproduce the model's creation process, including the training data.
In other words, while you can run and modify the model in your own environment, you cannot fully verify which data Cloudflare used to build it.
Fine-tuning begins as a joint effort with Cloudflare engineers
Cloudflare also announced a service for fine-tuning Clef to customer-specific operations.
In the first phase, Cloudflare's Forward-Deployed Engineer (FDE) team will work with customers to carry out fine-tuning. Based on that experience, the company plans to build a self-service system in which customers themselves can handle everything from data collection to training and redeployment.
In other words, although the models themselves are already public, a platform on which general users can freely start fine-tuning from a screen is not yet complete.
The vision Cloudflare presented combines several AI services it already offers.
Customers' own training datasets would be created from AI requests passing through AI Gateway, and Clef on Workers AI would be used to generate trial data. Containers would serve as isolated environments for reinforcement learning, used to score and replay AI agent behavior.
Using those results, a newly developed "Trainer" would update Clef's weights, and the result would ultimately be redeployed to Workers AI as a bring-your-own model.
Cloudflare says some of these pieces are still under development.
Rather than just providing a general-purpose decision model, the company appears to be aiming for a service that uses a company's actual decision data to turn Clef into a decision model dedicated to that company.
The current inquiry form also asks for the Cloudflare account ID and plan, as well as the specific use case and the AI Platform products in use.
At this point, it is better seen as a stage where Cloudflare builds the system together with customers who have concrete use cases, rather than a general service where anyone can choose a pricing plan and immediately start a training job.
Data handling differs between ordinary inference and fine-tuning
How customer data is handled also needs to be considered separately for ordinary inference and for fine-tuning.
Cloudflare says that in ordinary Clef usage it does not read or store requests and responses sent by customers, nor use them to train models.
The situation is different, however, when customers use the fine-tuning service.
In fine-tuning, the vision is to use customers' own requests and responses passing through AI Gateway as training datasets.
Companies that actually adopt it will therefore need to decide on operating policies, such as which communication logs to include in training and how to exclude personal and confidential information.
The next focus for Clef is not the short 38.8-millisecond response time itself.
What matters is how far misjudgments can be reduced, while keeping the required speed, when the model is given labeled data from real operations.
If that accuracy can be used to set rules such as "process automatically above this probability" or "send this kind of judgment back to a human," there may be more situations in which the fine-grained decisions now left to LLM-based AI agents can be replaced by faster decision models.
