watch Open-Weight Models · Research

Cloudflare Open-Sources Clef Decision Models as Ollama Adds a Decision-Model API

Data graphic: lollipop chart of median latency using Cloudflare's own figures: Clef-flash 38.8 ms, Clef 209.3 ms and Jev 524.1 ms.
PM

Supply chain security reporter · Updated Oct 2, 2026, 3:25 PM EDT

Cloudflare open-sourced Clef and Clef-flash, decision models that return probabilities in one pass, as Ollama v0.35 adds a /v1/systemone API.

Cloudflare released two open-weight "decision models", Clef (27B parameters) and Clef-flash (9B), on October 1, 2026. Three days earlier, Ollama v0.35.0 (September 28) added a /v1/systemone endpoint for the same class of model. The two moves target teams that want a fast, typed classifier inside an agent workflow instead of a free-text LLM call. No CVE or vulnerability is involved. This is a tooling story with a security-adjacent use case: Cloudflare says its Threat Intelligence team is testing Clef to classify website domains.

What a decision model is

Cloudflare describes a decision model as one that "produces bounded structured outputs cheaply, quickly and consistently" for a point in a workflow where a decision is needed. You pass a state (text, JSON, and for Clef also images or video) plus typed questions, and get probabilities back. Cloudflare's example is a support message: is it urgent, and which team should handle it. Your code then routes, escalates or defers to a human.

The Hugging Face cards say the model returns a probability for every allowed option of every question in a single forward pass, with no free-form generation and no output parsing. Cloudflare says the decision step is non-autoregressive, which is why it is faster than a generative LLM. The concept follows TypeSafe AI's Jev (System One), and both Cloudflare and Ollama describe their APIs as built on or compatible with it.

The API supports three question types: choice (probability per option), noul (probability a condition is true) and score (a score across ordered criteria).

The models

ClefClef-flash
Size27B9B
Base modelQwen/Qwen3.8-27BQwen/Qwen3.5-9B
LicenseApache-2.0Apache-2.0

Cloudflare says it froze each backbone and trained a joint routing head plus rank-256 low-rank adapters, with a Brier loss for calibration and a reinforcement-learning stage it calls RLCD. It claims a 64k context window against Jev's 32k and a vision encoder where Jev is text-only. These are vendor statements.

Vendor figures

All numbers below come from Cloudflare's announcement and its own Jev Decision Index runs. They have not been independently reproduced.

Vendor figureClefClef-flashJev
Median latency (ms)209.338.8524.1
p95 latency (ms)238.6122.4536.0
BFCL case exact98.4798.7695.75
BANKING77 macro-F194.2090.9379.74
CLINC150+OOS macro-F197.4366.7789.27
PhishNChips accuracy79.6075.0562.55
When2Call accuracy72.3765.5880.97

The table is mixed. Clef-flash trails sharply on CLINC150+OOS, and Jev wins When2Call. On PhishNChips, DiffusionGemma Jev scored 85.35, above both Clef models. Cloudflare says Clef leads the Jev Decision Index overall and that its models beat the other decision models on latency across 43 benchmarks, except Laya (5.8 ms median), which scored far lower on quality. In the Typesafe suite, Cloudflare reports Clef beating Jev in three of four areas shown; on security incidents the figures are 62.9 (Clef), 61.7 (Clef-flash) and 61.7 (Jev).

Cloudflare also says Clef fetched, rendered and classified a website in 2.2 seconds through its Browser Run tool, against 4.7 seconds for gpt-oss-120b, which returned only two classifications. That is a single vendor anecdote, not a benchmark.

Running them

Hosted. Both models run on Cloudflare Workers AI; Clef is addressed as @cf/cloudflare/clef. Cloudflare says it does not read, store or train on requests or responses to the hosted models.

Locally with Hugging Face. The weights are on Hugging Face (Cloudflare/clef, Cloudflare/clef-flash). The cards ship custom code (joint_schema_model.py) and were tested with torch 2.11 and transformers 5.10.2 on a single H200. The flow is snapshot_download, then load_release_model; a systemone helper accepts a Jev/SystemOne request body and returns the same response shape. A 27B model will need substantial GPU memory; the cards give no minimum.

In Ollama. Ollama v0.35.0 lists two decision models, Nimble (Bespoke Labs) and Tev1 (Together AI), pulled with ollama pull nimble. Neither the Ollama release notes, the Cloudflare blog nor the Hugging Face cards say Clef runs in Ollama, so do not assume it does. Both sides describe a Jev-compatible request shape, so client code written for one should be portable in principle, but that is untested here. Example against Ollama, from its release notes:

curl http://localhost:11434/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{"model":"nimble","state":"Our checkout has returned 500 errors since 9am.",
       "questions":{"label":{"type":"choice","instructions":"Which label fits this ticket?",
       "criteria":{"billing":"Payments and refunds","bug":"Software errors","account":"Login and account access"}}}}'

What defenders should do

  • Treat the figures as vendor claims and run your own labelled samples before putting a decision model in a triage or blocking path. Cloudflare's own table shows large per-task swings, including on phishing detection.
  • Use the returned probabilities and confidence for thresholds, with a human-review path for low-confidence results, as Cloudflare's own framing suggests.
  • Review the custom code in the Hugging Face repos before loading it, as with any model that ships Python.
  • Keep sensitive data in mind when choosing hosted versus local; the hosted data-handling statement is Cloudflare's own.

Cloudflare is also offering fine-tuning of Clef, first through a forward-deployed-engineer service and later a self-serve platform.

Sources

Keep reading

All latest →
  1. watchResearchXing4.0-29B-A4B: China Telecom's Open Coding Model Claims SWE-bench 75 on One GPU6 min
  2. watchResearchGitHub Copilot Drops Six Models on Oct 19: What Replaces GPT-5.5, GPT-5.4 and Grok 4.57 min
  3. elevatedResearchClaude Sonnet 4.5 Retires Nov 30: Six Requests That Return a 400 on Sonnet 5.58 min
  4. watchResearchGemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5: Is Google's New Model Actually Better?6 min
  5. stableResearchLing 3.1 Flash vs GLM 5.3 Flash vs Qwen 3.8 Flash Next: Three Bets on Cheap Agent Models6 min
  6. elevatedResearchLing-3.1-flash Is Ant Group's Best Flash Model Yet, and It Scores 87.9 on CyberGym6 min