clef flash benchmark

Clef-Flash Benchmarks: Latency, Cost and Memory Explained

Clef-Flash is reported at roughly 39 ms median latency, 122 ms p95 and $0.09 per million input tokens — here is what each published figure actually means.

The headline numbers

When Clef-Flash launched on 2026-10-01, the figures that travelled furthest were speed and price. Clef-Flash is a 9B parameter decision model post-trained from Qwen3.5-9B, and the published benchmark puts its median latency at around 39 ms with a p95 of roughly 122 ms.

Alongside latency, the reported input price sits near $0.09 per million tokens, and the model is said to run on roughly 41 GB of VRAM locally. It is also reported to be over ten times faster than Jev on the same decision tasks.

Those numbers describe different things — responsiveness, tail behaviour, cost and memory — and reading them as a single score is where most confusion starts. This guide separates them.

Latency: what median and p95 each tell you

The median, sometimes written p50, is the middle value: half of requests finish faster and half finish slower. A median near 39 ms for Clef-Flash means the typical decision returns in well under a tenth of a second, which is fast enough to sit inside an interactive request path.

The p95 is the value under which 95 percent of requests fall. A reported p95 near 122 ms means that one request in twenty takes longer than that. Tail latency is what users actually notice, so a p95 is usually the more honest number for capacity planning.

Reading both together tells you the shape of the distribution: a low median with a modest p95 suggests consistently quick responses with occasional slower ones, rather than a bimodal service that is sometimes instant and sometimes slow.

The Jev comparison

The published claim is that Clef-Flash is more than ten times faster than Jev on the same decision tasks. That comparison is only meaningful when the tasks are the same, which is why the wording matters.

A speed multiple is not a universal constant. It depends on the schema, the number of options per question, the length of the state and the hardware. Treat the multiple as a direction rather than a promise for your particular workload.

What the comparison does establish is intent: Clef-Flash was designed for latency-critical decisions, not for the heaviest reasoning, and its architecture and size reflect that priority.

Cost per million input tokens

The reported input price of about $0.09 per million tokens is the number that decides whether a hosted decision is cheap enough to run at volume. At that rate, a million calls that each consume a hundred tokens of state would cost roughly nine dollars in input.

That arithmetic is deliberately rough, because the true cost depends on how much state you send and how many questions you ask. A short ticket triage with a dozen options is far cheaper than a long document scored against a wide schema.

The useful habit is to estimate against your own traffic rather than against a benchmark article. Take your average tokens per call, multiply by your daily volume, and compare the result with the alternative you are currently paying for.

Memory: 41 GB compared with 85 GB

Clef-Flash needs roughly 41 GB of VRAM to run locally, while Clef needs about 85 GB. The gap is the practical consequence of the size difference between a 9B and a 27B model, and it changes what hardware is realistic.

About 41 GB fits comfortably on a single high-memory accelerator and can be squeezed onto well-equipped workstation cards with careful quantisation. Roughly 85 GB generally means multiple cards or a data-centre class machine, which is a different budget entirely.

This is why teams often start on Clef-Flash even when they intend to use Clef in production: the smaller model lets the whole pipeline be prototyped locally before any larger commitment.

Context window and vision

Both models accept a 64k-token context and include a built-in vision encoder, so the state can be long text, JSON, an image or even a video frame. That matters for benchmarks because the context length you actually use is one of the biggest levers on latency.

A decision that reads a short label is not the same workload as a decision that reads a full document and an attached image. The published latency figures describe the model, not every possible prompt you could send it.

Keeping the state as short as the decision allows is the simplest way to stay near the fast end of the reported distribution.

How to read a benchmark responsibly

Published figures come from the people who built the model, and independent measurements can differ. That is normal, and it is not a reason to distrust either source — it is a reason to check where a number came from before you plan around it.

Look at what was measured: which hardware, which tasks, which schema and which token counts. A benchmark on carefully chosen inputs can be accurate and still unrepresentative of a messy production stream.

The most useful benchmark is the one you run yourself, on your own states and your own schema, with the client you will actually deploy. Treat vendor figures as a starting point for that test rather than as its replacement.

Turning benchmarks into a workload estimate

Start with the decisions you already make. For each one, note how much state it reads, how many typed questions it answers and how many options each question allows. Options and questions are the parts that grow the scoring work.

Then measure locally or through the hosted API and compare the result with the published figures. If your numbers land close to the reported median and p95, the benchmark is representative for you. If they are far off, you have found the variable that matters most.

This is also how you decide between Clef-Flash and Clef. If Clef-Flash clears your latency budget with room to spare, there is rarely a reason to pay the larger model’s cost for that decision.

Reproducing the figures yourself

The quickest path is Ollama, which serves both clef-flash and clef as pulls. Running the 9B model locally lets you time real decisions without a network in the way, which isolates the model from the transport.

If local hardware is the constraint, the hosted API returns the same structured decision without the memory requirement, and the playground can show the output shape before you write any integration code.

Either route gives you the two numbers that matter most in practice: how long a typical decision takes and how long the slow ones take.

What stays stable as the numbers move

Benchmarks improve and prices change, but the interface does not. A decision is always a state plus a schema of typed questions returning a probability per allowed option, so an application written against that contract is insulated from the next round of performance news.

Choose the model for today’s numbers, but design the integration around the contract. When a faster or cheaper release appears, swapping the model behind the schema is a configuration change rather than a rewrite.

Try it here

Run this through the model

Paste your own state below and define a typed question. The frame is the live Clef-Flash demo; the panel on the playground page calls the API directly.

Clef-Flash community-hosted

Community-hosted demo of the 9B model. Runs live in your browser session.

Frequently asked questions

What is the median latency of Clef-Flash?

The published figure is around 39 ms median, with a p95 of roughly 122 ms on the decision tasks described at launch.

How much does Clef-Flash cost per million tokens?

The reported input price is near $0.09 per million tokens. Your real cost depends on how much state you send and how many typed questions you ask per call.

How much VRAM does Clef-Flash need?

About 41 GB to run locally, compared with roughly 85 GB for the larger Clef model.

Is Clef-Flash faster than Jev?

It is reported to be over ten times faster than Jev on the same decision tasks, though the exact multiple depends on your schema, state length and hardware.