clef vs clef flash

Clef vs Clef-Flash: choose precision or choose latency

Same family, same request shape, two very different budgets. Clef is the 27B model for hard decisions; Clef-Flash is the 9B model for decisions that must not slow a request down.

The two models side by side

Both models accept a state and a schema of typed questions and return probabilities. The difference is where they sit on the accuracy-versus-latency curve, and how much hardware they need.

  • Clef — 27B parameters, higher precision, roughly 85 GB VRAM locally.
  • Clef-Flash — 9B parameters post-trained from Qwen3.5-9B, about 39 ms median, roughly 41 GB VRAM locally.
  • Both — Apache 2.0, 64k context, vision encoder.

When the extra 18 billion parameters earn their keep

Reach for Clef when the decision is genuinely hard: ambiguous states, subtle distinctions, cases where a wrong answer is expensive and the latency budget is seconds rather than milliseconds. Batch scoring, offline review, contract triage and complex routing all fit.

Reach for Clef-Flash when the decision is in the request path: agent tool selection, live moderation, ticket routing, anything where 39 ms versus a few hundred milliseconds changes what you can build.

The honest default: prototype on Clef-Flash. Move the specific questions it gets wrong to Clef, rather than moving everything.

Running both

Ollama publishes both models, so a local switch is one pull away. On hosted infrastructure the two are separate endpoints with separate prices, and the larger model costs more per token because it is more expensive to serve.

Compare them on your own data

This site embeds a community playground for each model, on its own page. Running the same state through both, side by side, is the only benchmark that matters for your data.

Frequently asked questions

Which Clef model should I start with?

Start with Clef-Flash. It is faster, cheaper and sufficient for most bounded decisions. Move only the questions it fails to the 27B Clef model.

Do both models use the same API shape?

Yes. Both take a state plus typed questions and return probabilities per allowed option, so switching between them does not require a code rewrite.

What hardware do they need locally?

Clef-Flash needs about 41 GB of GPU VRAM and the 27B Clef model about 85 GB at single concurrency with a 64k context.

Keep reading

Model notes

One short email when a new decision model lands or a benchmark changes. No spam, unsubscribe in one click.