fine tune clef flash

How to Fine-Tune Clef-Flash on Your Own Decisions

Clef-Flash is open weights under Apache 2.0, so you can fine-tune the decision model itself to your vocabulary, your option sets and your calibration.

The weights are yours to adapt

Clef-Flash is released as open weights under the Apache 2.0 licence. That is a permissive licence: you can download the model, run it on your own hardware, modify it and use it commercially without asking permission or paying a royalty, provided you keep the licence notice intact. For a decision model, the practical consequence is that you are not borrowing a classifier that somebody else controls. You can own and adapt the exact model that scores your options.

Cloudflare publishes Clef-Flash as a 9B latency model post-trained from Qwen3.5-9B, and Clef as a 27B precision model. Both carry a 64k-token context window and a built-in vision encoder, and both are distributed through Workers AI, Ollama, Hugging Face and OpenRouter. The Hugging Face page is where the weights themselves live, and having the weights is what makes local adaptation possible in the first place.

What Apache 2.0 actually permits

Apache 2.0 grants a broad set of rights over the model weights: use, reproduction, modification, distribution and sublicensing, including inside commercial products. You can fine-tune the weights, merge them into an internal system or ship them as part of a paid service, as long as you preserve the attribution and licence notices and state any significant changes you made.

The licence also includes a patent grant from contributors, which matters when a model is embedded deeply in a commercial pipeline, and it is compatible with most corporate open-source policies. It does not grant you the Clef or Clef-Flash names, and it comes with no warranty. The licence gives you the building blocks, not a support contract.

For teams with data-residency or compliance constraints, this is often the deciding point. The weights can live inside your own perimeter, and a fine-tuned derivative can stay there too, which turns a hosted dependency into an internal asset.

The published reinforcement-learning path

Clef-Flash was not trained only by supervised next-token prediction. According to what Cloudflare published, it was post-trained from Qwen3.5-9B using reinforcement learning, meaning the model is optimized against a reward signal rather than merely imitating a corpus of correct answers.

That distinction matters enormously for a decision model, because the objective is not fluent prose. The objective is a well-calibrated probability distribution over a fixed answer space. A reinforcement loop can reward exactly that: did the model put most of its mass on the option that turned out to be right, and how confident was it when it was wrong?

You do not need to reproduce the exact recipe. Cloudflare published a path, not a hyperparameter sheet, and inventing numbers that were never released would help nobody. What you inherit is a model that already treats decisions as scored options, which gives your own fine-tuning a far better starting point than a general chat model would.

Why the objective is probabilities, not prose

When you fine-tune a chat model, success is often judged by how good the text looks. When you fine-tune a decision model, success is judged by whether the probabilities are useful. A confident wrong answer and a hesitant right one are different failure modes, and your evaluation should treat them differently.

Calibration is the property that makes a probability trustworthy: when the model says 0.8, the event should happen about eighty percent of the time. Fine-tuning is your opportunity to push calibration towards your domain instead of leaving it at some generic average learned from a much broader distribution of tasks.

Because every option is scored in a single pass and the answer space is typed, you can measure this directly. You compare the predicted distribution against observed outcomes and keep adjusting until the numbers behave the way your thresholds expect them to.

What fine-tuning can actually change

Fine-tuning is most useful for three things: domain vocabulary, the option sets that appear in your schema, and the calibration of the probabilities your application thresholds on. Those three cover most of the gap between a general decision model and one that fits a specific business.

Domain vocabulary covers the words and abbreviations your organisation uses that a general model has never seen in context: internal labels, product codenames, regulatory categories, ticket taxonomies. Adapting to that language reduces the ambiguity the model has to resolve at inference time and lets it spend its capacity on the decision rather than on decoding jargon.

Option sets and calibration are subtler. The schema stays flexible and defined at request time, but fine-tuning nudges the model towards the reasoning patterns behind your specific choices and towards the confidence levels your team has learned to trust from experience.

Data: how to collect decision examples

Fine-tuning a decision model needs examples of decisions, not essays. Each example is a state, the schema of typed questions you would send, and the correct answer for every question. The cleanest source is your own history: tickets that were triaged, content that was moderated, calls that were routed, applications that were reviewed.

The trick is to capture the state exactly as the model will see it at inference time and to record the outcome that was later confirmed, not just the first guess a human or a system made. Verified outcomes are worth far more than ambiguous ones, and a smaller set of clean labels usually beats a large set of noisy labels.

Where historical labels are thin, a careful human review pass can fill the gap, but consistency between reviewers matters more than raw volume. A decision model learns the pattern of your schema; contradictory labels simply teach it noise, and noise is expensive to unlearn.

Evaluation and calibration

Hold out a portion of your examples and never train on them. Then measure more than accuracy: look at the full distribution the model returns and compare it with what actually happened. Precision and recall on the top option are a start, but they hide the confidence information that makes a decision model valuable in the first place.

Calibration curves, where you bucket predictions by confidence and check the observed rate inside each bucket, show whether the model is overconfident or underconfident in your domain. If it is, that is a signal to keep tuning or to adjust your thresholds rather than to blindly trust the number.

Report the metrics that connect to your application: how often the top option is right, how often a low-confidence answer is escalated to a human, and how much the model agrees with your reviewers. Those are the numbers that justify shipping a fine-tune, and they are the numbers you will watch after deployment.

Running the adapted model locally

Clef-Flash needs roughly 41 GB of VRAM to run locally, and Clef needs about 85 GB. That puts a single fine-tuned Clef-Flash within reach of a workstation or a modest inference node, which is what makes the self-hosted path practical for a small team rather than only for large infrastructure groups.

You can pull the base model through Ollama, take the weights from Hugging Face, including quantised formats such as MLX 4-bit, and host it on your own infrastructure. The same model is also available through Workers AI if you would rather not manage hardware at all, which keeps the option open while you experiment.

After fine-tuning, the runtime footprint changes only modestly, because you are adapting an existing 9B model rather than growing it. The main cost is the training run itself, not a permanently larger serving cluster, and that is what makes iterative fine-tuning economically reasonable.

Staying compatible with the shared schema

One of the strongest reasons to fine-tune rather than replace is compatibility. Clef and Clef-Flash share an interface, and the schema is defined at request time rather than baked into the weights. A fine-tuned Clef-Flash can keep answering the same typed questions as the base model, which means your application code does not change when you swap the weights underneath it.

Keep the option sets and question types stable across training and serving. If you teach the model a new label, add it to the schema you send, and make sure your evaluation covers it. Compatibility is what lets you move a single decision between Clef-Flash, a fine-tune and the larger Clef model without rewriting the surrounding system.

Hosted or self-hosted: the decision

Hosted inference through Workers AI is the fastest way to start. There is no hardware to provision and no model to maintain, and you can delay the fine-tuning question until you have evidence that your schema and your data are worth adapting.

Self-hosting becomes attractive when you have strict data-residency requirements, a high and steady volume that makes per-token pricing expensive, or a genuine need for a model tuned to language no hosted endpoint has ever seen. The Apache 2.0 licence is what makes that option real rather than theoretical.

The two paths are not exclusive. Many teams prototype on the hosted endpoint, collect decision examples from real production traffic, and only then train and deploy a local fine-tune behind the exact same schema they were already using.

Try it here

Run this through the model

Paste your own state below and define a typed question. The frame is the live Clef-Flash demo; the panel on the playground page calls the API directly.

Clef-Flash community-hosted

Community-hosted demo of the 9B model. Runs live in your browser session.

Frequently asked questions

Can I fine-tune Clef-Flash commercially?

Yes. The weights are released under Apache 2.0, which permits modification and commercial use. You only need to preserve the attribution and licence notices and state any significant changes you made to the model.

Do I need reinforcement learning to fine-tune it myself?

No. Cloudflare published an RL post-training path for the original model, but your own adaptation can start with supervised examples of states, schemas and confirmed answers. RL is an option, not a requirement.

How much data do I need?

There is no published magic number. Start with a smaller set of clean, verified decision examples and grow it as you measure calibration, because consistent labels matter more than volume for a decision model.

Will a fine-tune still work with my schema?

Yes, if you keep the option sets and question types stable between training and serving. Clef-Flash defines its schema at request time, so the fine-tuned model answers the same typed questions and your application code does not change.