run clef flash locally

Run Clef-Flash locally

Clef-Flash is open-weight under Apache 2.0, so self-hosting is a legitimate option. Here is what it costs in hardware and how to get it running.

One command with Ollama

Ollama publishes the model, which makes the first run trivial. Pull it and send a state with a schema of typed questions; the model returns probabilities.

  • Install Ollama for your platform.
  • Run: ollama pull clef-flash
  • Send a state plus a JSON schema of typed questions and read the probabilities back.

Hardware you actually need

The published requirement is at least 41 GB of GPU VRAM for Clef-Flash at single concurrency with the full 64k context. The 27B Clef model needs about 85 GB. Apple silicon users can lean on community MLX 4-bit conversions on Hugging Face, which trade a little accuracy for a much smaller footprint.

If you only have a consumer GPU, quantised community builds or a hosted endpoint are the practical routes. Full precision wants a serious card.

When self-hosting makes sense

Self-host when the data cannot leave your infrastructure, when volume is high enough that per-token pricing dominates hardware cost, or when you need to fine-tune on your own labels. Cloudflare’s release includes reinforcement-learning fine-tuning, which is aimed squarely at that last case.

Host it when you need it working today, when traffic is spiky, or when the decision volume does not justify a GPU.

A middle path

You can keep the prototype on this site’s playground and move to a hosted endpoint when you go to production, without changing the request shape — the state-plus-schema contract is the same everywhere.

Frequently asked questions

Does Clef-Flash run on a laptop?

Not at full precision. It needs roughly 41 GB of GPU VRAM. Quantised community builds exist for Apple silicon and smaller cards, and hosted endpoints are cheaper than buying hardware for occasional use.

What is the Ollama command?

ollama pull clef-flash. The 27B variant is available as ollama pull clef.

Can I fine-tune it?

Yes. The weights are Apache 2.0 and Cloudflare published an RL fine-tuning path, so you can train it on your own decisions.

Keep reading

Model notes

One short email when a new decision model lands or a benchmark changes. No spam, unsubscribe in one click.