what is a decision model

What Is a Decision Model? Typed Probabilities Instead of Chat

A decision model takes a state and a schema of typed questions and returns a probability for every allowed option in one pass — no chat, no prose.

A decision model in one sentence

A decision model is a machine learning model that reads a state — the situation it must reason about — together with a schema of typed questions, and returns a probability for every allowed option in a single pass. It does not write prose and it does not chat. It scores.

Clef and Clef-Flash are open-source examples of this design, released by Cloudflare under the Apache 2.0 licence. Clef is a 27B precision model and Clef-Flash is a 9B latency model post-trained from Qwen3.5-9B. Both accept a 64k-token context and both ship with a built-in vision encoder.

The result of a call is structured: for each typed question you define, you receive one probability per allowed answer. That structure is what separates a decision model from a general-purpose language model, and it is why the output can be consumed by ordinary application code.

Why the chat interface gets in the way

When you ask a chatbot to decide something, it answers in sentences. The probability is buried in adjectives, the options are implied rather than enumerated, and the shape of the output changes every time you ask.

Developers then write parsers, regular expressions and heuristics to recover a label or a number from that text. Every repair is a guess about what the model meant, and every guess can be wrong in a way that is expensive to detect.

A decision model removes that whole layer. If your schema says the answer must be one of approved, rejected or needs-review, then those three options are the only answers the model can return, each with a probability attached.

The anatomy of a decision

Every decision has three parts. The state is the raw material: a support ticket, a document, a frame from a video, a row of telemetry. The schema is the set of typed questions you want answered about that state. The output is the probability the model assigns to each allowed option.

The state can be text, JSON or an image. Clef and Clef-Flash handle all three because their vision encoder is part of the base model rather than a bolt-on component that only some tasks can use.

The schema is where the engineering happens. You are not asking a vague question and hoping for a sensible reply; you are defining the exact answer space the model is permitted to work with.

What typed questions actually are

A typed question constrains the answer. An enumeration question lists the exact labels the model may choose from. A boolean question is a two-way split. A bounded number question gives the model a range and expects a value inside it.

Because the options are typed, the model can never invent a fourth category or return a number outside the range. Anything that is not in the schema is not a valid decision, and the interface makes that explicit.

This constraint is a feature, not a limitation. It means downstream systems can trust the shape of the answer before they even look at the probabilities, which simplifies validation enormously.

One pass, every probability

A decision model scores all allowed options at once in a single forward pass, rather than generating one token at a time. That design keeps latency low and makes the probabilities comparable, because every option is scored from the same evaluation of the same state.

Clef-Flash was built for latency-critical work. Published figures put its median latency at around 39 ms and its p95 at roughly 122 ms, with input pricing reported near $0.09 per million tokens. Clef targets higher precision at a larger size.

Because the options are scored together, the model is also internally consistent in a way that a model answering several open questions in sequence cannot be, since a free-form answer can quietly contradict an earlier one.

Where Clef and Clef-Flash sit in the family

Clef is the precision model: 27B parameters, roughly 85 GB of VRAM to run locally, and the higher accuracy you would expect from its size. Clef-Flash is the latency model: 9B parameters, about 41 GB of VRAM, and reported to be over ten times faster than Jev on the same decision tasks.

The two share the same interface, so a schema written for one works on the other. You can prototype on Clef-Flash and escalate to Clef when a particular decision needs more care, without rewriting the application around it.

Both are released under Apache 2.0, which matters for teams that need to run decisions inside their own perimeter rather than sending every state to a hosted endpoint.

How it differs from a classifier

A classical classifier returns one label from a fixed list that was baked in during training. A decision model returns probabilities for several typed questions at once, and the questions can change from request to request without any retraining.

That flexibility is the difference between maintaining a bespoke model for every task and running one model that serves many tasks, each described in the schema you send with the request.

The probability output also suits thresholds better than a hard label does, because you decide where to draw the line rather than accepting whatever single category the model happened to pick.

How it differs from a general LLM

A general LLM can answer almost anything in almost any format, which is exactly why it is hard to depend on for a decision. Its confidence is usually unstated and difficult to calibrate against real outcomes.

A decision model trades breadth for reliability. It does one job — scoring a fixed answer space — and returns a number you can threshold, log, route on and audit after the fact.

The two are complementary rather than competing. A general LLM can summarise or explain; a decision model can make the call.

What you can build with it

Common patterns include support triage with fields for category, urgency and escalation; content moderation with allow, review and block; agent tool routing that chooses which function to call; and structured extraction into a fixed form.

In each case the state changes constantly but the schema stays stable, which is the sweet spot for a decision model. The model does not need to learn your business logic; it needs to score the options your logic already defined.

  • Support triage: category, urgency, escalation.
  • Moderation: allow, review or block.
  • Agent routing: which tool or function to call next.
  • Structured extraction into a fixed schema.

Try a decision model now

The fastest way to understand the idea is to use one. The playground lets you write a state, add typed questions and see the probability for every allowed option, all in one screen and with no setup.

If you prefer to read first, the API documentation covers the request and response shape in detail, and the pricing page explains what a hosted call costs before you commit to anything.

Try it here

Run this through the model

Paste your own state below and define a typed question. The frame is the live Clef-Flash demo; the panel on the playground page calls the API directly.

Clef-Flash community-hosted

Community-hosted demo of the 9B model. Runs live in your browser session.

Frequently asked questions

Is a decision model the same thing as a chatbot?

No. A chatbot generates prose and leaves the probability implicit. A decision model scores a fixed set of allowed answers and returns a probability for each one in a single pass.

Do I have to define the allowed options myself?

Yes. The schema of typed questions is yours to write. The model cannot return an answer outside the options you listed, which is what makes the output safe to consume in code.

What is the difference between Clef and Clef-Flash?

Clef is the 27B precision model; Clef-Flash is the 9B latency model. They share an interface, so the same schema works on both and you can move between them freely.

Are the models open source?

Both are released under the Apache 2.0 licence and are distributed through Workers AI, Ollama, Hugging Face and OpenRouter.