Clef-Flash, the 9B decision model that answers instead of chatting
Clef-Flash turns a state and a schema of typed questions into decisions — a probability for every allowed option, in a single forward pass. Around 39 ms median, Apache 2.0 weights, 64k context, vision built in.
What Clef-Flash actually is
Clef-Flash is the smaller of two decision models released by Cloudflare on 1 October 2026. It is a 9-billion-parameter multimodal model post-trained from Qwen3.5-9B, and it is not a chatbot. Given a state — text, JSON, an image, or video frames — and a schema of typed questions, it returns a probability for each permitted answer.
That shape is the whole point. A classifier wants a label; a router wants a destination; a reviewer wants a verdict with a confidence score. Clef-Flash returns exactly that, in one pass, without a second parsing step and without a prompt that tries to talk a language model into behaving like a classifier.
- 9B parameters, vision encoder included.
- 64k token context window.
- Apache 2.0 licence, open weights.
- Typed answers with probabilities, one forward pass.
How fast, and what that costs
The published benchmark numbers put Clef-Flash at roughly 39 ms median latency and 122 ms at p95, which makes it more than ten times faster than Jev on the same decision tasks. Input tokens run about $0.09 per million when you call it through a hosted provider.
The speed is not a marketing footnote. It is what moves a decision model out of an offline batch job and into the request path — an agent routing tool calls, a moderation gate, a support ticket triage step that has to finish before a human sees the queue.
Where you can run it
Clef-Flash is available on Cloudflare Workers AI, through Ollama with a single pull, on Hugging Face including community MLX 4-bit conversions, and on OpenRouter. Local inference needs at least 41 GB of GPU VRAM; the larger 27B Clef model needs about 85 GB.
- Hosted: Workers AI, OpenRouter.
- Local: ollama pull clef-flash.
- Apple silicon: MLX 4-bit community builds on Hugging Face.
- This site: an embedded community playground plus a hosted API.
Try it before you read another paragraph
The playground on this site embeds a live community demo of Clef-Flash so you can paste a state and define questions in the browser. Nothing is installed, nothing is stored, and the model itself runs elsewhere.
If you want the same behaviour in your own stack, the API page shows the request shape and the pricing page shows the hosted tiers.
Frequently asked questions
Is Clef-Flash free?
The weights are Apache 2.0, so the model itself is free to download and self-host. Hosted inference is billed by whichever provider you use. This site lets you try it in the browser for free.
How is Clef-Flash different from a normal LLM?
A language model generates text and you parse it. Clef-Flash returns a probability for every allowed option directly, so there is nothing to parse and no way for it to answer outside the schema you gave it.
How much VRAM does Clef-Flash need?
About 41 GB for local inference at single concurrency with a 64k context. The larger 27B Clef model needs roughly 85 GB. If that is out of reach, use a hosted endpoint.
What is the difference between Clef and Clef-Flash?
Clef is the 27B precision model; Clef-Flash is the 9B latency model. Choose Clef when accuracy on hard decisions matters more than speed, and Clef-Flash when the decision sits in a request path.
Keep reading
Model notes
One short email when a new decision model lands or a benchmark changes. No spam, unsubscribe in one click.