Cloudflare has released Clef, a family of open-weight models built to choose between predefined options instead of generating text. The release, made during Birthday Week, includes 27B and 9B multimodal versions plus a platform for adapting them to decision tasks. The move matters because routing, escalation and triage inside AI agents could shift from slow text generation to fast classification.

Cloudflare releases Clef decision models for AI agents

Clef models, latency figures and compatibility

The larger Clef model takes a state and a schema of typed questions as input and returns a probability for each allowed option for every question. It accepts text, JSON, images or video and produces typed outcomes in a single forward pass, with no free-form text and no output parsing. Cloudflare positions this output as direct input for agent logic, for example routing a support request, escalating it, or deferring the decision to a human.

The design centers on placement in the agent hot path. Michelle Chen, Alex Reneau and Kevin Flansburg say hosting on Cloudflare infrastructure uses GPUs at the edge, which cuts network latency and speeds up decisions. The stated combination is Clef for the decision step together with an LLM on Workers AI for the action step. Cloudflare also says its API is compatible with the Typesafe AI Jev System One model.

Latency is the sharpest contrast between the two sizes. Clef-Flash, the 9B model for latency-sensitive decisions, shows median latency of 38.8 ms in Cloudflare benchmarks, against 209.3 ms for the 27B Clef model. The larger model adds a vision encoder for classifying visual content, while Jev handles only text classification today, according to the authors. Clef also offers a 64k context window versus 32k for Jev, allowing more input state per classification.

What fast classification means for agent builders

For companies running support triage, bot detection or request routing, a classifier with millisecond response changes where automated judgment can sit. Instead of calling a general text model and parsing its answer, teams get structured probabilities that plug into rules, thresholds and escalation paths. Smaller firms gain a downloadable starting point, while larger operations can standardize routing behavior across high-volume queues without adding parsing code.

The limits sit around accuracy, calibration and fit to local data. Hacker News commenters questioned whether decision models form a new category and warned that public benchmarks are easy to overfit. A Reddit user put the test differently: check whether Clef knows when to defer, especially where a false positive costs more than a miss, and see if confidence drops when data shifts. The benchmark figures alone do not answer that.

The marker to watch is Cloudflare fine-tuning for support triage and bot classification on historical labeled data, plus customer adaptation through its fine-tuning service. That service starts with help from Cloudflare engineers, with self-service planned later and no firm date given. If tuned versions ship on Workers AI and Hugging Face feedback shows stable deferral behavior, classification as a separate agent layer will look confirmed for business use.