What agent tool routing really is
An agent decides which tool to call next. It reads the current state of the task, considers the functions it has been given, and picks one. That step is called tool routing, and it happens many times inside a single run.
Most agents implement routing as a chat completion. The tools are described in a prompt, the model is asked to pick one, and the reply is parsed back into a function name. It works often enough to ship, and it fails in ways that are annoying to debug.
A decision model treats routing as what it is: a typed question with a fixed set of allowed answers. The tools become the options, and the model returns a probability for each one.
Why chat-based routing is fragile
When routing goes through a chat completion, the tool name arrives inside prose or a half-formed JSON blob. If the model invents a function that does not exist, the agent crashes or silently falls through to a default.
The bigger problem is that the choice is unquantified. You get a name without a confidence, so you cannot tell a near-tie between two tools from a clear winner, and you have nowhere to attach a fallback rule.
Every fix becomes a prompt edit, and prompt edits are guesses. The routing logic ends up spread across examples, instructions and parsers instead of living in code you can review.
The schema of a routing decision
With a decision model, the state is the conversation so far, the task goal and any intermediate results. The schema is an enumeration question listing the exact tools the agent may call.
Clef and Clef-Flash read that state and schema together and return one probability per tool in a single pass. Clef is the 27B precision model; Clef-Flash is the 9B latency model post-trained from Qwen3.5-9B. Both accept a 64k-token context and include a vision encoder.
The agent never sees a tool name that was not in the enumeration. The routing decision is constrained before the model runs, which removes a whole class of runtime errors.
The enumeration is the entire contract. If a tool does not exist yet, it does not appear in the schema, and the agent cannot call it. When a new tool ships, you add one option and canary it by rolling out the schema change like any other release.
Scoring candidates instead of generating text
Because every candidate is scored from the same evaluation of the same state, the probabilities are directly comparable. A score of 0.72 for search and 0.19 for database is meaningful in a way that two separate prose answers never are.
That comparability lets you add a margin rule: only call a tool when its probability clears a threshold and leads the runner-up by a set margin. Otherwise the agent can ask a clarifying question or stop.
It also lets you log the full distribution. After an incident you can see whether the model was torn between two tools or confident and simply wrong, which is a different problem with a different fix.
Arguments are a separate question
Choosing a tool is only half the decision. Once a tool is selected, the agent needs arguments, and those can also be expressed as typed questions in the same request.
A bounded number question can pin down a page size or a limit. An enumeration can fix a sort order. A free-text field can capture a search phrase when the state demands it.
Running selection and slot-filling in one pass keeps the agent loop tight and avoids a second round trip that would otherwise add latency to every step.
When no tool should be called
A reliable router must be able to say no. Some turns need clarification, some need a final answer, and some need the agent to stop entirely.
Adding candidates such as ask-clarification or finish as options in the enumeration gives the model a legitimate way to decline the obvious tool. Without that escape hatch, a forced choice produces confident nonsense.
This single change fixes a large share of runaway agents, because the model no longer has to invent a call when none is warranted.
Latency inside the agent loop
Routing happens on every step, so its latency multiplies across a run. A slow router makes the whole agent feel sluggish even if each individual tool is fast.
Clef-Flash was built for exactly this shape of work. Published figures put its median latency around 39 ms and its p95 near 122 ms, with input pricing reported close to $0.09 per million tokens.
It runs on about 41 GB of VRAM locally, while Clef needs roughly 85 GB, so a latency-sensitive loop fits on a single accelerator while higher-precision runs can escalate to the larger model without changing the schema.
Caching helps too. Routing requests for identical states are cacheable, and the deterministic shape of the output makes that safe in a way that prose responses are not.
Guardrails and auditability
A constrained router is easier to guard. Because the output space is fixed, you can reject any response outside it before it reaches your executor, and you can cap the number of times a given tool may run.
Probabilities also feed policy. If the top tool is destructive and its score is not decisive, the agent can defer to a human instead of acting. That rule is ordinary application code, not another prompt.
Every routing decision can be stored with its distribution, its threshold and the tool chosen, giving you an audit trail that explains agent behaviour after the fact.
A worked pattern
Picture a research agent with three tools: search, fetch-page and summarise. Each turn it receives the goal and what it has gathered so far.
The router scores the three tools plus finish. Search wins early, fetch-page dominates once a promising URL exists, and finish rises when the gathered context already answers the goal. All of it is one request per turn with a distribution you can threshold.
- State: goal, history, intermediate results.
- Schema: search, fetch-page, summarise, finish.
- Rule: act only when the leader clears a threshold and a margin.
- Log: the full distribution per turn.
Route your first agent call
The quickest way to test the idea is to describe a state, list the tools as options and read the probabilities. The playground does exactly that with no setup and no code.
The API documentation covers the request and response shape for wiring it into a real loop, and the article on decision models explains why typed probabilities beat prose for this kind of choice.