Cloudflare Adds Its First Team-Trained Decision Models to Workers AI
Clef and Clef-flash return probabilities for predefined answers. Their weights are available under the Apache 2.0 license, alongside hands-on fine-tuning support.

Cloudflare added its first Workers AI team-trained decision models to the platform on Oct. 1, 2026, giving software agents tools to make structured choices such as classifying a request or deciding whether to escalate it.
The models, Clef and Clef-flash, return probabilities for predefined answers rather than generating free-form text. Their weights are available under the Apache 2.0 license, and Cloudflare introduced a hands-on reinforcement-learning fine-tuning service for customers adapting Clef to their workloads.
Clef has 27 billion parameters and a 64,000-token context window. Clef-flash has 9 billion parameters and the same context capacity. Clef accepts text, JSON, images or video; requests can include up to four images.
Median latency across 43 benchmark runs was 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash. Typesafe’s Jev had a median latency of 524.1 milliseconds. On the BANKING77 benchmark, macro-F1 scores were 94.20 for Clef, 90.93 for Clef-flash and 79.74 for Jev. Clef scored highest on seven of the 10 decision benchmarks in the company’s tests.
The models run on GPUs across Cloudflare’s network, close to users, for decisions within agent workflows. The format lets an application select from defined outcomes, such as a request category, before another model or system takes action.
Clef follows the System One API, so developers with an existing Jev integration can switch by changing the endpoint and model. Cloudflare is initially offering the fine-tuning service with hands-on support.