The headline numbers
When Clef-Flash launched on 2026-10-01, the figures that travelled furthest were speed and price. Clef-Flash is a 9B parameter decision model post-trained from Qwen3.5-9B, and the published benchmark puts its median latency at around 39 ms with a p95 of roughly 122 ms.
Alongside latency, the reported input price sits near $0.09 per million tokens, and the model is said to run on roughly 41 GB of VRAM locally. It is also reported to be over ten times faster than Jev on the same decision tasks.
Those numbers describe different things — responsiveness, tail behaviour, cost and memory — and reading them as a single score is where most confusion starts. This guide separates them.
Latency: what median and p95 each tell you
The median, sometimes written p50, is the middle value: half of requests finish faster and half finish slower. A median near 39 ms for Clef-Flash means the typical decision returns in well under a tenth of a second, which is fast enough to sit inside an interactive request path.
The p95 is the value under which 95 percent of requests fall. A reported p95 near 122 ms means that one request in twenty takes longer than that. Tail latency is what users actually notice, so a p95 is usually the more honest number for capacity planning.
Reading both together tells you the shape of the distribution: a low median with a modest p95 suggests consistently quick responses with occasional slower ones, rather than a bimodal service that is sometimes instant and sometimes slow.
The Jev comparison
The published claim is that Clef-Flash is more than ten times faster than Jev on the same decision tasks. That comparison is only meaningful when the tasks are the same, which is why the wording matters.
A speed multiple is not a universal constant. It depends on the schema, the number of options per question, the length of the state and the hardware. Treat the multiple as a direction rather than a promise for your particular workload.
What the comparison does establish is intent: Clef-Flash was designed for latency-critical decisions, not for the heaviest reasoning, and its architecture and size reflect that priority.
Cost per million input tokens
The reported input price of about $0.09 per million tokens is the number that decides whether a hosted decision is cheap enough to run at volume. At that rate, a million calls that each consume a hundred tokens of state would cost roughly nine dollars in input.
That arithmetic is deliberately rough, because the true cost depends on how much state you send and how many questions you ask. A short ticket triage with a dozen options is far cheaper than a long document scored against a wide schema.
The useful habit is to estimate against your own traffic rather than against a benchmark article. Take your average tokens per call, multiply by your daily volume, and compare the result with the alternative you are currently paying for.
Memory: 41 GB compared with 85 GB
Clef-Flash needs roughly 41 GB of VRAM to run locally, while Clef needs about 85 GB. The gap is the practical consequence of the size difference between a 9B and a 27B model, and it changes what hardware is realistic.
About 41 GB fits comfortably on a single high-memory accelerator and can be squeezed onto well-equipped workstation cards with careful quantisation. Roughly 85 GB generally means multiple cards or a data-centre class machine, which is a different budget entirely.
This is why teams often start on Clef-Flash even when they intend to use Clef in production: the smaller model lets the whole pipeline be prototyped locally before any larger commitment.
Context window and vision
Both models accept a 64k-token context and include a built-in vision encoder, so the state can be long text, JSON, an image or even a video frame. That matters for benchmarks because the context length you actually use is one of the biggest levers on latency.
A decision that reads a short label is not the same workload as a decision that reads a full document and an attached image. The published latency figures describe the model, not every possible prompt you could send it.
Keeping the state as short as the decision allows is the simplest way to stay near the fast end of the reported distribution.
How to read a benchmark responsibly
Published figures come from the people who built the model, and independent measurements can differ. That is normal, and it is not a reason to distrust either source — it is a reason to check where a number came from before you plan around it.
Look at what was measured: which hardware, which tasks, which schema and which token counts. A benchmark on carefully chosen inputs can be accurate and still unrepresentative of a messy production stream.
The most useful benchmark is the one you run yourself, on your own states and your own schema, with the client you will actually deploy. Treat vendor figures as a starting point for that test rather than as its replacement.
Turning benchmarks into a workload estimate
Start with the decisions you already make. For each one, note how much state it reads, how many typed questions it answers and how many options each question allows. Options and questions are the parts that grow the scoring work.
Then measure locally or through the hosted API and compare the result with the published figures. If your numbers land close to the reported median and p95, the benchmark is representative for you. If they are far off, you have found the variable that matters most.
This is also how you decide between Clef-Flash and Clef. If Clef-Flash clears your latency budget with room to spare, there is rarely a reason to pay the larger model’s cost for that decision.
Reproducing the figures yourself
The quickest path is Ollama, which serves both clef-flash and clef as pulls. Running the 9B model locally lets you time real decisions without a network in the way, which isolates the model from the transport.
If local hardware is the constraint, the hosted API returns the same structured decision without the memory requirement, and the playground can show the output shape before you write any integration code.
Either route gives you the two numbers that matter most in practice: how long a typical decision takes and how long the slow ones take.
What stays stable as the numbers move
Benchmarks improve and prices change, but the interface does not. A decision is always a state plus a schema of typed questions returning a probability per allowed option, so an application written against that contract is insulated from the next round of performance news.
Choose the model for today’s numbers, but design the integration around the contract. When a faster or cheaper release appears, swapping the model behind the schema is a configuration change rather than a rewrite.