Cloudflare’s open decision models, runnable in your browser

Turn any state into a decision — in one pass

Clef-Flash is a 9B multimodal decision model. Give it a situation and a schema of typed questions; it returns a probability for every allowed answer. No chat, no prompt engineering, around 39 ms median.

39 ms median latency, published9B parameters, vision included64k token context windowApache 2.0 open weights
Clef-Flash playground community-hosted

Community-hosted demo of the 9B model. Runs live in your browser session.

How it works

How a decision model works

Three steps, one request. If your prompt ends with “answer with one of these options”, this is the model that was built for it.

01

Describe the state

Paste a support ticket, a JSON payload, a product description or an image reference. The state is the situation the model has to reason about — up to 64k tokens.

02

List the answers you allow

Write typed questions, each with a closed set of options. The schema is the contract: the model cannot answer outside it, so there is nothing to parse and no way for it to improvise.

03

Read the distribution

Every question comes back with a probability per option and a winning answer. Threshold on confidence, log the rest, and stop writing regexes against generated prose.

Comparison

Decision model vs LLM

Same input, different contract. This is the distinction that decides your architecture.

Large language modelClef-Flash
OutputFree text you have to parseA probability per allowed option
ShapeOpen-ended; the format driftsClosed set, guaranteed by the schema
Cost of a mistakeA broken parser in productionA low confidence score you can threshold on
Best atWriting, explaining, extractingRouting, labelling, scoring, moderating
Use cases

What people use it for

Anywhere a system has to choose between bounded options before a human sees the queue.

Agent tool routing

Pick the next tool from a fixed catalogue before the agent spends a token on deliberation.

Support triage

Label intent and urgency, then route on the confidence score rather than on a guess.

Content moderation

Return a category and a probability you can defend after the fact.

Image labelling

The vision encoder takes an image as the state, so the same schema works for pictures.

Lead scoring

Bounded bands instead of a number a language model invented this time.

Quality gates

Approve, reject or escalate — with the distribution logged for review.

The facts, before the pitch

  • Two models: Clef 27B for precision, Clef-Flash 9B for latency. Both Apache 2.0 with open weights.
  • Clef-Flash is post-trained from Qwen3.5-9B with a vision encoder and a 64k context window.
  • Published latency for Clef-Flash is about 39 ms median and 122 ms at p95 — over ten times faster than Jev on comparable tasks.
  • Local inference needs roughly 41 GB of GPU VRAM for Clef-Flash and about 85 GB for the 27B Clef model. Both are on Ollama.
  • Available on Cloudflare Workers AI, Ollama, Hugging Face and OpenRouter.

Questions people ask first

Is this site Cloudflare?

No. It is an independent playground. Clef and Clef-Flash are open-source models from Cloudflare, released under the Apache 2.0 licence, and this site is not affiliated with them.

What is the difference between Clef and Clef-Flash?

Clef is the 27B precision model; Clef-Flash is the 9B latency model. Start with Clef-Flash, and move only the questions it gets wrong to the larger model.

Do I need an account to try it?

No. The browser playground works without signing up. An API key only raises the daily decision quota.

Do you store the text I paste?

We store the count, the model and the latency — not the state itself. The privacy policy lists exactly what is kept.

Open the playground

Run a state through Clef-Flash right now. Nothing to install, no account needed.

Model notes

One short email when a new decision model lands or a benchmark changes. No spam, unsubscribe in one click.