decision model vs classifier

Decision Model vs Classifier: Fixed Labels or Typed Probabilities

A classifier learns one fixed set of labels during training. A decision model reads your schema at request time and scores every allowed option in one pass.

The short answer

A classifier and a decision model can both return a label, but they get there in completely different ways. A classifier learns a fixed mapping between inputs and a set of categories during training. A decision model reads your schema at request time and scores every allowed option in that schema.

The practical consequence is retraining. When you need a new category, a classifier needs fresh labelled data and another training run. A decision model needs a different schema in the request body. That single difference reshapes how quickly a team can react to a new requirement.

The distinction is not about accuracy, either. A well-trained classifier can be very accurate on the tasks it was built for. The question is what happens when the task shifts: whether you can change behaviour without collecting data, retraining a model and redeploying it.

What a classical classifier actually does

A classical classifier takes an input and returns one label from a set that was fixed during training. The label space lives inside the weights, so adding a category means changing the model itself rather than the request.

It also returns confidence in a form that is easy to print but hard to use. A single softmax value tells you how strongly the model preferred its choice, not how likely that choice is to be correct in the world.

This design is efficient and well understood. Its limits only start to matter when the problem keeps changing, or when you need more than one decision from the same input.

What a decision model does instead

Clef and Clef-Flash, the open-source decision models released by Cloudflare under the Apache 2.0 licence, take a different route. They read a state — a ticket, a document, a frame of video — together with a schema of typed questions, and return a probability for every allowed option in a single pass.

Clef is the 27B precision model and Clef-Flash is the 9B latency model post-trained from Qwen3.5-9B. Both accept a 64k-token context and both include a vision encoder, so the same interface handles text, JSON and images.

Because the schema is part of the request, nothing about the model has to change when your categories change. You send a new schema and the answer space moves with it.

Why retraining is the real cost

The obvious cost of a classifier is the compute to train it. The hidden cost is everything around that: collecting labelled examples, keeping them balanced, waiting for a training run, and validating that the new model did not regress on the old categories.

Every one of those steps takes people and calendar time. A team that ships a new moderation category once a quarter is not making a technical choice; it is responding to how expensive the change has become.

A decision model moves that work into the schema. Adding an option or a question becomes a code change, reviewable like any other, and it can ship the same day.

Where the probability changes your code

A hard label forces a binary decision at the model boundary. A probability lets you move that decision into your application, where you can see it, log it and tune it.

If a support ticket scores 0.91 for billing and 0.86 for escalation, you can route on the first and flag the second. With a single label you would have had to accept one answer and lose the tension between them.

Thresholds also make the system auditable. You can raise or lower a cutoff without touching the model, and you can explain to a reviewer exactly why a case was routed a certain way.

Several questions, one pass

A classifier answers one question. A decision model answers as many typed questions as you put in the schema, scoring all of them together from the same reading of the state.

That shared pass matters for consistency. When a model answers questions in sequence, a later answer can quietly contradict an earlier one. When every option is scored from one evaluation, the answers are drawn from a single interpretation of the input.

It is also cheaper to operate than running several classifiers side by side, because one request replaces many.

For an agent or a workflow, that single pass also removes a coordination problem. There is no need to reconcile the outputs of several models or to decide which one to trust when they disagree, because the disagreement is already expressed as a probability inside one response.

When a classifier is still the right tool

None of this makes classifiers obsolete. If your categories are genuinely fixed, your volume is enormous and your latency budget is tiny, a small purpose-built classifier can be the most efficient option available.

The decision model shines when the answer space changes, when you need several related decisions at once, or when you need calibrated probabilities instead of a single argmax label.

The honest framing is that they sit at different points on a trade-off. A classifier is narrow and cheap; a decision model is flexible and still fast enough for interactive work.

Latency, cost and deployment

Speed is usually the first objection raised. Clef-Flash was built for latency-critical paths: published figures put its median latency around 39 ms and its p95 near 122 ms, with input pricing reported close to $0.09 per million tokens.

It runs on about 41 GB of VRAM locally, while Clef needs roughly 85 GB, so the latency model comfortably fits on a single modern accelerator. Both are distributed through Workers AI, Ollama, Hugging Face and OpenRouter.

That range of distribution means a team can start on a hosted endpoint and later move the same schema to a local deployment, or the reverse, without touching the application logic.

Migrating without a rewrite

If you already run a classifier, the migration is smaller than it looks. Keep the categories you have, express them as an enumeration question, and add the extra questions you always wanted but could not justify training a second model for.

You can run both systems side by side, compare their outputs on real traffic, and move the threshold rather than flipping a switch. Because the decision model needs no training, there is no waiting period before you can evaluate it.

Teams often keep the classifier as a fallback for a while and retire it once the probability-based routing is demonstrably better on the cases that matter.

Try both shapes

The clearest way to see the difference is to hold the same input against both. The playground lets you write a state, define typed questions and read the probability for every option without writing any code first.

From there, the API documentation describes the request and response shape, and the pricing page explains what a hosted call costs before you commit to anything.

Try it here

Run this through the model

Paste your own state below and define a typed question. The frame is the live Clef-Flash demo; the panel on the playground page calls the API directly.

Clef-Flash community-hosted

Community-hosted demo of the 9B model. Runs live in your browser session.

Frequently asked questions

Can a decision model replace my classifier entirely?

Often, yes. If your categories are stable and the schema maps cleanly onto typed questions, a decision model can sub in without retraining. High-volume, very cheap classification can still be served by a small dedicated classifier.

Do I need to retrain a decision model when categories change?

No. The schema of typed questions is sent with each request, so changing the allowed options is a request change rather than a training job.

Is a decision model slower than a classifier?

Clef-Flash is built for latency-critical work, with published median latency near 39 ms and p95 near 122 ms, which is fast enough for interactive traffic as well as batch jobs.

What can a decision model do that a classifier cannot?

It scores several typed questions at once and returns a probability for every option, so you can threshold, route and audit instead of accepting one hard label.