# Cloudflare’s Clef Model Posts Less Than Half Jev’s Median Latency

By Simon Yoon

Canonical URL: https://www.tokenpost.com/news/technology/26982
Published: 2026-10-06T08:03:58.000Z
Updated: 2026-10-06T08:03:58.000Z
Section: Technology

> Clef recorded 209.3 milliseconds across 43 benchmarks, compared with 524.1 milliseconds for TypeSafe’s Jev. Hosted input pricing for Clef is about 5.7 times higher.

Cloudflare’s Clef model recorded a median latency of 209.3 milliseconds across 43 benchmarks, less than half TypeSafe’s Jev median of 524.1 milliseconds, while hosted Clef input costs about 5.7 times more per token.

Cloudflare charges $0.24 per million input tokens for Clef and $0.09 for Clef-flash. Jev costs $0.042 per million input tokens, with no charge for output tokens. The models return probabilities for predefined choices rather than open-ended text.

Released Oct. 1, Clef and the smaller Clef-flash are the first machine-learning models trained by Cloudflare’s Workers AI team. Both are available through Workers AI, with weights offered under the Apache 2.0 license for download and self-hosting. This provides an alternative to hosted inference pricing.

Clef is based on Qwen3.8-27B, while Clef-flash uses Qwen3.5-9B. Their API is compatible with Jev’s API, also known as System One.

In 10 quality benchmarks, Clef or Clef-flash led seven, Jev led two, and a model based on Google’s DiffusionGemma led the remaining test. On BANKING77, a customer-service intent classification test, Clef scored 94.20 compared with Jev’s 79.74. Jev scored higher than Clef on When2Call and BRIGHT.

Latency varied by model and measurement. Clef-flash’s median was 38.8 milliseconds. At the 95th percentile, latency was 238.6 milliseconds for Clef, 122.4 milliseconds for Clef-flash and 536.0 milliseconds for Jev.

Cloudflare is also providing hands-on support for reinforcement-learning fine-tuning, a process for adapting a model using feedback. A self-service version is planned for later.
