What multimodal means here
In this context multimodal is a narrow, useful claim. It means the state a decision model reasons over can be an image or a video frame, not only text or JSON. It does not mean the model captions the picture, chats about it or writes a description. It scores the visual input against your schema and returns a probability for every option you allowed.
Clef and Clef-Flash both ship with a built-in vision encoder, released on 2026-10-01 under the Apache 2.0 licence with a 64k-token context. Clef is the 27B precision model and Clef-Flash is the 9B latency model post-trained from Qwen3.5-9B. Neither needs a separate vision service standing beside it.
So when a photo of a damaged parcel arrives, the model does not tell you what it sees. It answers the questions you asked — is it damaged, how severe, which region — each with a probability attached, in a single call.
The vision encoder is part of the base model
This is the detail that changes how you deploy the system. The vision encoder is not an adapter bolted onto a text-only checkpoint, and it is not a separate model that converts an image into words before the real model starts work. It is part of the base model, trained alongside the reasoning path.
Because it is built in, image states and text states travel through the same weights, the same interface and the same deployment. You do not maintain a second endpoint for visual inputs or a routing layer that decides which model to call based on input type.
That keeps the operational surface small. One model, one schema format, one place to look when something behaves differently than you expected.
States that are images or video frames
The state is the raw material of a decision. For a text workflow it might be a support ticket or a log line. For a visual workflow it is a photograph, a screenshot, a scan, a product image or a single frame extracted from a video stream.
Video is simply a sequence of states. You sample frames at whatever interval your problem needs, then score each frame against the same schema. A ten-minute clip at one frame per second is six hundred decisions, each one structured and each one comparable to the last.
The state does not have to be a tidy studio image. Screenshots of interfaces, dashboard captures and camera frames all work, because the encoder learns to read the signal rather than expecting a particular framing.
A schema of typed questions applies to visual states
The schema does not change just because the state is visual. A typed question still constrains the answer. An enumeration question lists the exact labels the model may choose from. A boolean question is a two-way split. A bounded number question expects a value inside a range you define.
For a photo of a retail shelf you might ask whether the product is in stock, how many facings are visible within a range, and which category the shelf belongs to. Each of those is a typed question with a finite answer space, and the model returns one probability per allowed option.
The constraint is what makes the result safe to consume. The model cannot invent a fourth category or return a count outside the range, so downstream code can trust the shape of the answer before it inspects a single probability.
Where visual decisions earn their keep
Visual question answering rarely needs prose. It needs a label, a routing decision or a threshold, and that is exactly what a decision model produces. The same call can drive an alert, a queue assignment or a human review flag.
Moderation is the obvious case: a frame is allowed, sent to review or blocked, with the probability telling you how close the call was. Inspection is the industrial version, where a photo of a part is scored for defect type and severity without a caption ever being written.
Frame triage is where throughput matters most. A long recording contains mostly uneventful frames, so you score every sampled frame cheaply and keep only the ones that cross a threshold for deeper attention.
- Visual QA routing: send an interface or product shot to the right queue.
- Moderation: allow, review or block a frame.
- Inspection: defect type and severity from a single photo.
- Frame triage: keep only the moments that cross a threshold.
Why one-pass scoring matters for video throughput
Video multiplies everything. A pipeline that scores one image per request can afford a leisurely model, but a pipeline that scores hundreds of frames per stream cannot. Per-frame latency compounds directly into how far behind real time the stream falls.
A decision model scores all allowed options at once in a single forward pass rather than generating one token at a time. Clef-Flash was built for this kind of work. Published figures put its median latency around 39 ms and its p95 near 122 ms, with input pricing reported near $0.09 per million tokens and performance reported at over ten times faster than Jev on comparable tasks.
Because the option set is scored from one evaluation of one frame, the probabilities stay comparable frame to frame. That consistency is what lets you set a global threshold instead of recalibrating for every new image that arrives.
64k context and multiple frames
Both models accept a 64k-token context, which is enough to carry several frames alongside text instructions in the same request. You can score a short sequence rather than an isolated still, which helps when the decision depends on movement or change between frames.
There is a real trade-off here. Images consume tokens, so more frames or higher resolution means fewer of both. You decide how to spend the budget: a few detailed frames, or many coarse ones across a wider window.
A practical pattern is to sample frames deliberately, group a small number per request, and keep the schema identical across the group so the answers line up when you aggregate them.
The same schema across text and image states
One of the quieter advantages of a built-in encoder is that text and image states share a schema. A moderation rubric written for text carries over to a screenshot without being rewritten, and the same questions run against both kinds of input.
This lets you unify what would otherwise be two pipelines. A support flow can read a written complaint and a screenshot of the failing screen with the same typed questions, then merge the probabilities into one routing decision.
It also means no retraining when a new input type appears. You describe the new state in the request and reuse the schema you already have, so the application logic does not fork just because the input changed shape.
How it differs from captioning and vision chat
A captioning model turns an image into a sentence. That sentence is pleasant to read and nearly impossible to threshold. The signal you care about is buried in word choice, and the shape of the output changes with every image.
A vision chat model answers open questions in prose. It is flexible, but the answers are hard to compare, hard to audit and easy to contradict. Neither approach gives application code a fixed answer space it can rely on.
A multimodal decision model commits to that answer space up front. It scores the frame against your typed questions and returns a number for every option, which is the difference between a description of an image and a decision about one.
How to try it
The fastest way to see this working is to use it. The playground lets you supply a state, add typed questions and read the probability for every allowed option on one screen, with no setup and no account gymnastics.
If you would rather run it yourself, Clef-Flash is available through Ollama with a single pull command, and the documentation covers the request and response shape for both models. Workers AI, Hugging Face and OpenRouter are the other distribution paths.
Start with one visual question you genuinely need answered, define its options narrowly, and watch how much simpler the result is than anything a caption would give you.