Guides
Plain-language guides to decision models: how they differ from LLMs, how fast Clef-Flash is, and where a bounded answer beats a paragraph.
Multimodal Decision Models: Deciding on Images and Video
A multimodal decision model scores an image or a video frame against a schema of typed questions and returns a probability for every allowed option in one pass.
9 min · 2026-10-14
How to Fine-Tune Clef-Flash on Your Own Decisions
Clef-Flash is open weights under Apache 2.0, so you can fine-tune the decision model itself to your vocabulary, your option sets and your calibration.
10 min · 2026-10-13
Clef-Flash vs Jev: Two Decision Models Compared
Clef-Flash and Jev share the same decision-model interface, so the real choice is a trade-off between latency-critical and precision-critical work.
9 min · 2026-10-12
Support Ticket Triage with AI: Category, Urgency and Escalation at Once
Support ticket triage is three decisions, not one. A decision model returns a probability for category, urgency and escalation together in a single pass.
9 min · 2026-10-11
AI Moderation API: Probabilities Instead of a Black-Box Verdict
An AI moderation API should return probabilities per label, not a single verdict, so your team owns the threshold that decides what gets blocked.
9 min · 2026-10-10
Agent Tool Routing with a Decision Model: Score Every Function
Agent tool routing is a decision, not a conversation. A decision model scores every candidate function in one pass and returns a probability per tool.
9 min · 2026-10-09
Decision Model vs Classifier: Fixed Labels or Typed Probabilities
A classifier learns one fixed set of labels during training. A decision model reads your schema at request time and scores every allowed option in one pass.
9 min · 2026-10-08
How to Run Clef-Flash on Ollama
Clef-Flash ships as an Ollama model: pull it, write a schema of typed questions and read a probability for every allowed option without leaving your machine.
8 min · 2026-10-07
Clef-Flash Benchmarks: Latency, Cost and Memory Explained
Clef-Flash is reported at roughly 39 ms median latency, 122 ms p95 and $0.09 per million input tokens — here is what each published figure actually means.
9 min · 2026-10-06
What Is a Decision Model? Typed Probabilities Instead of Chat
A decision model takes a state and a schema of typed questions and returns a probability for every allowed option in one pass — no chat, no prose.
8 min · 2026-10-05