Cloudflare Clef is a decision model: you give it an input and a list of typed questions, and it returns a probability for each allowed answer rather than writing a sentence back. It ships as two open-weight models, Clef and the smaller Clef-flash, both hosted on Workers AI and released on 1 October 2026. The job is narrow: it classifies, so your own code can decide what happens next.
Cloudflare documents it on the Workers AI model page and the Hugging Face model card, which both say there is no free-form text generation and nothing to parse out of the output.
How it differs from a chatbot
A large language model is general-purpose. It is most often used to generate prose or code, and it can also be constrained to emit structured output, such as JSON that matches a schema. Clef is narrower: it uses a decision head that scores every option you allowed in one forward pass and produces no free-form text, reading the probabilities straight from the backbone's internal representations. It still has an LLM inside: Cloudflare post-trained it from a Qwen backbone and added a small scoring head, so a language model sits underneath, aimed at one task. And whichever model you use, it is the application or harness around it that executes a tool call or an action, not the model itself; Clef returns only the decision.
The three question types
You describe each decision as a typed question and can ask up to 64 in one request. Clef uses the same types as the System One API it is compatible with (Cloudflare changelog):
noulis a yes/no question. It returns the probability that the answer is yes, between 0 and 1. The probability is the answer; there is no separate confidence number here.choicepicks one option from a set you define. It returns the chosen option, a probability for each option, and a confidence value. The chosen option is the highest-probability one; the confidence is a separate value worth reading on its own rather than assuming it equals the top probability.scorerates the input against an ordered rubric you supply, such as "No impact, Minor, Major, Critical". It returns a probability per level and a probability-weighted score, which can land between the named levels rather than snapping to one. That score is the model's weighted estimate, not an objective measurement.
A support-triage example
Cloudflare's own example is a support queue. A ticket says checkout has been failing for every customer for the last hour; you send that text as the state and three questions: is it urgent (noul), which team should own it, billing, technical, or sales (choice), and how severe is the impact (score). Clef returns a probability for urgency, a chosen team with a probability for each option, and a weighted severity score.
Clef makes the call, but it does not act on it. Your code reads the numbers and decides, routing on a probability threshold you set, paging someone when the severity score passes a threshold you have validated, and sending anything ambiguous to a person. A refund request can be flagged, but the refund should run through your own authorization rules rather than fire because a probability crossed a line. The model infers and the application executes. (Any thresholds here are illustrative; Clef ships no universal cutoff.)
Clef, Clef-flash, and open weights
Clef is the larger 27B model Cloudflare positions for highest-precision decisions; Clef-flash is a 9B model for latency-critical, hot-path calls. Both are post-trained from Qwen backbones, 27B for Clef and 9B for Clef-flash, and both carry a vision encoder, so they can read images as part of the state. Both are released under the Apache 2.0 license with open weights you can download and run yourself. Open weights do not mean a free hosted service: running Clef on Workers AI is a paid API, while the open-weights route means hosting it yourself.
How much input it reads
The model page lists a context window of 65,536 tokens for both models. The hosted API notes that a long input state is truncated to fit the model's token limit, so check how your chosen endpoint behaves on input the size you actually send. The open-weights helper is a separate matter: its encode step defaults to 16,384 tokens, a helper default that differs from the model's capacity rather than a hard ceiling. On other inputs, the hosted API documents up to four embedded images per request, with size caps and no image URLs. Video is supported in the open weights, but there is no documented hosted video parameter, so do not assume video works through the hosted API.
Clef and Jev
Clef is API-compatible with Jev, TypeSafe's first System One model, and uses the same noul, choice, and score question types (TypeSafe's docs), so Cloudflare says you can move an existing Jev integration to Clef by changing the endpoint and the model name. One clear difference is licensing: Clef and Clef-flash are released as open weights under Apache 2.0, while Jev is TypeSafe's proprietary model. Compatible is not identical, though: two different models can give different probabilities, confidence values, and calibration on the same input, so validate a migration on your own data rather than trust it blind.
What it does not promise
A probability is a claim from a model, not a measured fact, and Clef can be wrong: it can misclassify, be pushed by adversarial or unusual input, and carry bias from its training data. A returned confidence of 0.9 is not a guarantee the label is right nine times in ten; Cloudflare trained for calibration with a Brier-score objective and a reinforcement-learning step, but calibration on their data does not transfer automatically to yours. So the advice is ordinary: set action thresholds from your own validation set, keep a person in the loop for high-stakes actions such as refunds, account changes, or security blocks, and keep permissions in your application. Read the benchmarks as vendor-reported, workload-specific numbers rather than independent results, and test on cases like yours. The fine-tuning service that sharpens it for one workload is at the hands-on design-partner stage today, not self-serve for everyone.
Frequently asked questions
What is Cloudflare Clef in simple terms?
A decision model from Cloudflare. You give it an input and a set of typed questions, and it returns a probability for each allowed answer instead of writing text back, so your own code can route, escalate, or defer. It comes as two open-weight models, Clef and the smaller Clef-flash, hosted on Workers AI and released on 1 October 2026.
How is Clef different from a chatbot or an LLM?
A general chatbot is usually used to generate prose or code, and it can also be constrained to emit structured JSON. Clef runs a single forward pass and returns probabilities over the options you defined, with no free-form text. It still has a language model inside, a Qwen backbone that Cloudflare post-trained, but it is pointed at classification rather than writing, and the application around it is what acts on the decision.
Can Clef replace a human in decisions?
Not on its own. Clef classifies and returns probabilities; your application decides what to do with them and holds the authorization for any action. It can misclassify and its confidence is not a guarantee of accuracy, so keep human review for high-stakes actions and set your thresholds on your own data.
Sources
- Cloudflare blog, Introducing Clef: our open-source decision models, and new RL fine-tuning platform, 1 October 2026.
- Cloudflare Workers AI docs, clef and clef-flash (parameters, context window, input limits).
- Hugging Face model cards, Cloudflare/clef and Cloudflare/clef-flash (backbone, outputs, Apache 2.0 license).
- Cloudflare changelog, AI product group (question types and API).
- TypeSafe, System One documentation and typesafe.ai (Jev and the shared question types).