Why run Clef-Flash locally
Clef-Flash is a 9B decision model, and that size is what makes local deployment practical. The published VRAM requirement is around 41 GB, which fits on a single high-memory accelerator rather than a fleet of cards.
Running locally keeps every state on your own machine, which matters when the inputs are customer tickets, medical notes or internal documents that you would rather not send to a hosted endpoint.
It also removes the network from the latency picture. The reported median of roughly 39 ms describes the model itself, and a local run lets you see that number without transport variation on top.
What you need before you start
You need a machine with enough accelerator memory — around 41 GB for Clef-Flash — and a recent install of Ollama. A CUDA or Metal-capable GPU is strongly preferred, because a decision model on CPU will be dramatically slower than the published figures.
You do not need a framework, a Python environment or a training pipeline. The model arrives quantised and ready to serve, and the rest of the work is writing a schema and calling the local endpoint.
If your hardware falls short, you can still follow along using the hosted API or the browser playground, both of which return the same structured decision.
- Ollama installed and running.
- About 41 GB of accelerator memory for Clef-Flash.
- A little disk space for the model weights.
- A state and a schema of typed questions to test with.
Pulling the model
The model is distributed through the Ollama library, and Clef-Flash pulls under the name clef-flash. The larger 27B precision model is available as clef on the same library, so you can hold both alongside each other.
The first pull downloads the weights and can take a while depending on your connection. Subsequent runs start from the local cache, so the download is a one-time cost rather than a per-session one.
Once the pull completes, the model is available to any tool that speaks to the local Ollama endpoint, including the command line and the HTTP API.
Making your first decision
A decision request has two inputs: the state and the schema. The state is the situation — a ticket, a message, a document or an image. The schema lists the typed questions you want answered and the allowed options for each.
Unlike a chat completion, you are not asking for a paragraph. You are asking the model to score each allowed option, so the response you get back is a probability per option, ready to threshold or log directly.
Start with something small. A single enumeration question with three or four options is enough to confirm the model is responding and that your client is reading the structure correctly.
Writing a schema that works
A good schema names each question clearly and lists only the options you would genuinely accept. Enumeration questions cover categorical choices, boolean questions cover binary splits, and bounded number questions cover values within a range.
Keep the questions independent where you can. If two questions overlap, the model may return probabilities that are individually sensible but jointly inconsistent, and you will spend time reconciling them downstream.
The options should be mutually exclusive and exhaustive. If a real answer could fall outside your list, add an option for it rather than forcing the model to choose the least-wrong label.
Tuning for latency
The biggest lever on local latency is the state length. Trim the state to what the decision actually needs; a long preamble that does not change the answer costs time on every call.
The second lever is the size of the schema. Each additional question and each additional option adds scoring work, so a lean schema is a fast schema. Ask only the questions whose answers you will use.
After that, standard local inference practices apply: keep the model resident rather than reloading it, avoid swapping, and measure your own p95 rather than trusting a single timing.
Running Clef alongside Clef-Flash
Both models live in the same library, so you can pull clef as well and compare them on the same schema. Because the interface is identical, the only change in your request is the model name.
The 27B Clef model needs roughly 85 GB of VRAM, so it may not fit on the same machine as Clef-Flash. Running it on a stronger box and pointing your client at that endpoint keeps the comparison honest.
A common pattern is to route easy decisions to Clef-Flash and reserve Clef for the cases where a small difference in probability changes the outcome. Both return the same shape of answer, so the routing logic stays simple.
Using the hosted API instead
If local hardware is unavailable, the hosted API exposes the same decision contract without the memory requirement. That is the fastest way to get an integration working before you invest in a machine.
The API is also the practical choice for bursty traffic, since you pay for calls rather than for idle accelerator time. The reported input price is near $0.09 per million tokens, which makes a hosted prototype inexpensive.
You can move between local and hosted without changing your schema, because the request and response shapes are the same. Prototyping on one and shipping on the other is a supported path.
Troubleshooting the common issues
If the model refuses to load, the usual cause is insufficient accelerator memory. Check the reported requirement for Clef-Flash — about 41 GB — and confirm nothing else is holding the device.
If responses feel slower than the published median, look at state length and schema size first, then at whether the model is being reloaded between calls. Fixing the reload alone can transform the timings.
If the answers look wrong rather than slow, the schema is usually the problem. Loosen an option, split an overlapping question or add a catch-all, and the probabilities tend to become far easier to interpret.
Where to go next
Once a local decision works, the next step is to measure it on a realistic sample. Collect a few hundred real states, run them through the same schema and look at how the probabilities distribute across the options.
From there, choose your thresholds, decide whether Clef-Flash is accurate enough for the task, and integrate the call into your application. The decision itself is only the middle of the pipeline; the surrounding logic is what turns a probability into an action.
The playground, the API documentation and the model pages all describe the same contract, so anything you learn in one transfers directly to the others.