The Clef model family at a glance
Two open-weight decision models, one request shape. This page is the reference card: parameters, context, latency, hardware and distribution.
The table that answers most questions
Everything below is what Cloudflare and the community published around the 1 October 2026 release. Numbers are quoted as published rather than re-measured here.
- Clef — 27B parameters, precision model, ~85 GB VRAM locally.
- Clef-Flash — 9B parameters, post-trained from Qwen3.5-9B, ~39 ms median, ~$0.09 per million input tokens, ~41 GB VRAM locally.
- Both — Apache 2.0, 64k context, vision encoder, typed answers with probabilities.
Why “decision model” is a separate category
The family is trained to produce a distribution over a fixed answer set rather than to continue text. That is why a decision model can be small and fast while still being useful: it is not asked to know everything, only to choose well among the options you allow.
Where to get them
Workers AI hosts both. Ollama distributes both. Hugging Face carries the weights, including community quantisations. OpenRouter exposes the smaller model through its unified API.
Which one to open first
If you are evaluating, run Clef-Flash first — it is the one you are most likely to ship. Use the playground pages here to compare it against the 27B model on the same state.
Frequently asked questions
Are the Clef models free?
The weights are released under the Apache 2.0 licence, so you can download, self-host and fine-tune them at no licence cost. Hosted inference is billed by the provider.
Which model is faster?
Clef-Flash. The published median is about 39 ms, more than ten times faster than Jev on comparable decision tasks.
Do the models take images?
Yes. Both ship with a vision encoder, so a state can be text, JSON, an image or video frames.
Keep reading
Model notes
One short email when a new decision model lands or a benchmark changes. No spam, unsubscribe in one click.