Cloudflare
Clef and Clef-flash
Cloudflare's first trained models: open decision models that answer typed questions with probabilities

Key facts
- 1 Oct 2026Workers AI and Hugging Face
- Released
- 27B / 9BClef / Clef-flash
- Sizes
- $0.24 / $0.09per M input tokens
- Price
- 65,536tokens on Workers AI
- Context
- 209 / 39 msCloudflare's own runs
- Median latency
- Apache 2.0open weights
- Licence
Clef is a decision model from Cloudflare: it reads a block of text, JSON, images or video plus a list of typed questions, and returns a probability for every allowed answer, so software can act on the numbers. Cloudflare released it on 1 October 2026 in two sizes, Clef with 27 billion parameters and Clef-flash with 9 billion, hosted on Workers AI at $0.24 and $0.09 per million input tokens and published as open weights under Apache 2.0. They are the first models Cloudflare's Workers AI team has trained itself.
Clef is a decision model: an AI model that reads a block of input, such as a support ticket, a web page or a screenshot, and answers a list of fixed questions with a probability for every allowed answer. Cloudflare released it on 1 October 2026 in two sizes, Clef with 27 billion parameters and Clef-flash with 9 billion, hosted on its Workers AI platform and published as open weights. They are, in Cloudflare’s words, its “first Cloudflare-trained ML model from the Workers AI team”.
Decision models are a new category. TypeSafe’s Jev opened it in September, and Clef answers the same request format, so a program built for Jev can switch by changing the endpoint and the model name. The name comes from music: a clef tells a reader what each line on a stave means, “and the CF hearkens to Cloudflare”, the company says.
You ask it typed questions and get probabilities back
Clef takes a state and up to 64 questions in one request and returns a probability for every allowed answer. Each question is one of three types: a yes-or-no question that returns the probability of yes, a choice from a set you define, or a score against an ordered rubric, the Workers AI changelog says. A program then routes a ticket, triggers an escalation or hands the case to a person on the numbers.
The state can be plain text, JSON, images or video. Clef carries a vision encoder, which Cloudflare names as a difference from Jev, and accepts up to four embedded PNG, JPEG or WebP images a request. Cloudflare has been testing it with its Threat Intelligence team to classify website domains: sorting one site into categories took Clef 2.2 seconds to fetch, render and classify, against 4.7 seconds for the company’s fastest general model, gpt-oss-120b, which also returned fewer classifications.

Two sizes, both built on Qwen
Clef is post-trained from Alibaba’s Qwen3.8-27B and Clef-flash from Qwen3.5-9B, with the base weights frozen. Cloudflare trained rank-256 low-rank adapters and a routing head on top, using a loss that rewards valid answers and a Brier score that rewards honest probabilities, on its own synthetic data. At inference Clef runs one pass over the input and then scores every valid answer in parallel, so there is no text to generate token by token.
| Clef | Clef-flash | |
|---|---|---|
| Base model | Qwen3.8-27B | Qwen3.5-9B |
| Parameters, Hugging Face | 27,356,728,560 | 9,409,813,744 |
| Context on Workers AI | 65,536 tokens | 65,536 tokens |
| Input | Text, JSON, images, video | Text, JSON, images, video |
| Questions per request | 1 to 64 | 1 to 64 |
| Licence | Apache 2.0 | Apache 2.0 |
The context window is 64K tokens, which Cloudflare sets against Jev’s 32K. The open reference code on Hugging Face defaults to 16,384 tokens, which a user can raise.
Clef costs $0.24 per million input tokens
Workers AI charges $0.24 per million input tokens for Clef and $0.09 for Clef-flash, on Cloudflare’s pricing page read on 2 October 2026. Cloudflare lists an input price only. Against the other decision models, both cost more per token:
| Decision model | Maker | Price per million input tokens | Where |
|---|---|---|---|
| Clef | Cloudflare | $0.24 | Workers AI |
| Clef-flash | Cloudflare | $0.09 | Workers AI |
| Jev 1.13 | TypeSafe | $0.042 | OpenRouter |
| pplx-decider-v1-27b | Perplexity | $0.04 | Perplexity API |
| Strands Decider 2B | AWS Strands Labs | Free download | Your own hardware |
Cloudflare also offers fine-tuning. Customers work with its forward-deployed engineers first, and a self-serve platform for training and redeploying the model on Cloudflare follows later, the company says. Cloudflare promises that it does not read, store or train on requests or responses unless a customer opts into fine-tuning.
How Clef scores against Jev
On ten decision benchmarks Cloudflare selected and ran itself, a Clef model scores highest on seven and Jev on two, while a DiffusionGemma-based Jev model tops the tenth, PhishNChips, at 85.35. The figures are Cloudflare’s own run of the community Decision Index 0.2.1 suite.
| Benchmark | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL, case exact | 98.47 | 98.76 | 95.75 |
| ToolRet, nDCG@10 | 69.19 | 66.43 | 65.28 |
| API-Bank, accuracy | 91.93 | 93.11 | 88.19 |
| Home appliances, case exact | 82.95 | 97.73 | 52.27 |
| When2Call, accuracy | 72.37 | 65.58 | 80.97 |
| BANKING77, macro-F1 | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS, macro-F1 | 97.43 | 66.77 | 89.27 |
| BRIGHT, nDCG@10 | 45.91 | 39.26 | 47.52 |
| Amazon ESCI, macro-F1 | 57.48 | 57.39 | 55.21 |
| PhishNChips, accuracy | 79.60 | 75.05 | 62.55 |
On the combined leaderboard Cloudflare publishes, Clef sits first at 61.21, ahead of Jev at 57.91 and Clef-flash at 57.07, on a chance-corrected scale where 0 is random guessing. The board labels both Clef rows “Self-reported” and says they “have not been reproduced by the upstream board”. Jev keeps the lead on the hardest general tests in the full table on Hugging Face: 78.3 on GPQA Diamond against 48.0 for Clef, and 82.7 on MMLU-Pro against 65.9.

An outside board puts Jev ahead. On Benchmark Heaven’s JevBench v1.5.5, Jev 1.13 scores 80.0 on capability against 77.3 for Clef and 70.3 for Clef-flash, and 88 on calibration against 86.7 and 87.6. Its official ranking weighs intelligence, calibration, speed and cost equally, with the Clef models’ cost estimated for self-hosting, and puts Clef-flash 25th and Clef 62nd of 109 systems; priced at Workers AI’s rates, they would rank 24th and 43rd.
Clef-flash is the fast one
Across Cloudflare’s 43 benchmark runs, Clef-flash answered in a median 38.8 milliseconds and Clef in 209.3, against 524.1 for Jev’s hosted API. Cloudflare’s board notes that the Jev figure is a network round trip from its lab and “Not comparable to the on-card single-process figures”.
| Latency, Cloudflare’s runs | Clef | Clef-flash | Jev |
|---|---|---|---|
| Median | 209.3 ms | 38.8 ms | 524.1 ms |
| 95th percentile | 238.6 ms | 122.4 ms | 536.0 ms |
The Jev figure depends on where it is measured. YFarmX timed Jev at a median 230 milliseconds over 31 calls on 18 September 2026, on a three-question routing task. On that figure Clef runs about level with Jev, and Clef-flash is the clearly faster model.
Where to run it
Clef runs through the Workers AI binding in a Cloudflare Worker, through the REST API, or behind Cloudflare’s AI Gateway, with the model IDs @cf/cloudflare/clef and @cf/cloudflare/clef-flash. The weights and reference code sit on Hugging Face under Apache 2.0, the licence of the Qwen base models, for anyone who wants to run them on their own GPUs.
Cloudflare now trains its own models
Until 1 October, Workers AI hosted other labs’ models. Clef is the first one Cloudflare built, and it launched with an open licence, a price list and a fine-tuning offer on the same day. The things to watch are the self-serve fine-tuning platform Cloudflare has promised, and whether the upstream Decision Index board reproduces Clef’s self-reported lead. For the model that defined the category, see Jev; for the other launches of the same week, Perplexity’s pplx-decider-v1-27b and Strands Decider 2B.
Questions people ask
- What is Cloudflare Clef?
- Clef is a decision model that Cloudflare released on 1 October 2026. It reads an input such as a support ticket, a web page or an image, plus up to 64 typed questions, and returns a probability for every allowed answer: yes or no, one option from a set, or a level on a rubric. It comes in two sizes, Clef with 27 billion parameters and Clef-flash with 9 billion, and they are the first models Cloudflare's Workers AI team has trained.
- How much does Clef cost?
- On Workers AI, Clef costs $0.24 per million input tokens and Clef-flash $0.09, according to Cloudflare's pricing page read on 2 October 2026. Cloudflare lists an input price only. The open weights are free to download from Hugging Face under the Apache 2.0 licence.
- Is Clef better than TypeSafe's Jev?
- It depends on the board. On Cloudflare's own run of the community Decision Index 0.2.1, Clef scores 61.21 against Jev's 57.91, and a Clef model is top on seven of ten benchmarks Cloudflare selected. The leaderboard marks Clef's scores as self-reported. On Benchmark Heaven's JevBench v1.5.5, Jev 1.13 leads both Clef models on its capability score, 80.0 against 77.3 for Clef and 70.3 for Clef-flash.
- Can I run Clef on my own hardware?
- Yes. Cloudflare published both models on Hugging Face under Apache 2.0 with reference code. Clef is post-trained from Qwen3.8-27B and Clef-flash from Qwen3.5-9B, so Clef needs a GPU with room for a 27-billion-parameter model in BF16. Hosted use on Workers AI takes the model ID @cf/cloudflare/clef or @cf/cloudflare/clef-flash.
More in Large Language Models
All LLMs →- TypeSafeJeva decision model that returns a typed answer and a confidence
- PerplexityPerplexity Decider v1 27Bthe decision model behind Perplexity's Decisions API, with open weights
- AWS Strands LabsStrands Decider 2Ba small open decision model that runs on a gaming GPU or a laptop
- AlibabaQwen3.8-Flash-Nextthe cheap tier, and a preview of what Qwen4 is built on
- Google DeepMindGemini 4 Argonlonger reasoning, coding and cyber defence
- OpenAIGPT-5.6the flagship from 9 July to 3 September 2026