support ticket triage ai

Support Ticket Triage with AI: Category, Urgency and Escalation at Once

Support ticket triage is three decisions, not one. A decision model returns a probability for category, urgency and escalation together in a single pass.

What triage actually decides

Every incoming ticket forces the same set of small decisions. Which queue does it belong to. How urgent is it. Does it need an escalation. Who should own it.

Those decisions are made thousands of times a day, often in seconds, and they set the tone for everything that follows. Route a billing question to a technical queue and the customer waits twice as long for an answer.

Support ticket triage with AI is the practice of making those decisions consistently and immediately, before a human has read the message.

The cost of getting it wrong

Misrouting is expensive in a way that is easy to underestimate. It burns the customer's time, the agent's time and the trust that a fast first response buys.

It also compounds. A ticket that bounces between queues collects internal notes, loses its context and often arrives at the right person with less patience than it started with.

And it is invisible at scale. Nobody measures the tickets that were routed slightly wrong and still resolved, so the slow leakage of quality never shows up in a dashboard.

There is a customer-experience cost too. A ticket routed to the wrong specialist is often answered by someone who has to transfer it, and every transfer is a moment where a customer wonders whether anyone is actually listening.

Triage as a decision model problem

A decision model reads a state — here, the ticket text, its metadata and the history around it — together with a schema of typed questions, and returns a probability for every allowed option in one pass. It does not write a summary and it does not chat.

Clef and Clef-Flash are the open-source decision models behind this approach, released by Cloudflare under Apache 2.0. Clef is the 27B precision model; Clef-Flash is the 9B latency model post-trained from Qwen3.5-9B. Both accept a 64k-token context and include a vision encoder.

That means a screenshot attached to a ticket can be part of the same decision as the text, without a separate pipeline.

Category, urgency and escalation in one pass

The schema for triage usually holds three enumerations: a queue or category, an urgency band, and an escalation flag. You can add a language, a product area or a sentiment question if you need them.

Instead of three separate calls, the model scores all of them together from one reading of the ticket. The answers arrive consistent with each other, which matters when urgency depends on category.

A billing question about a failed payment is not the same as a billing question about an invoice format, and the probabilities let you express that difference rather than collapsing it into one label.

The order matters less than the coherence. Because urgency is scored alongside category, a ticket in a quiet queue can still register high urgency when the wording demands it, instead of being forced into a single global judgement.

  • Category: billing, technical, account, other.
  • Urgency: low, normal, high.
  • Escalation: no, yes.
  • Optional: language, product area, sentiment.

Reading the whole thread, not the subject line

Subject lines are noisy and customers describe symptoms rather than causes. A ticket titled "it is down" may be a local network problem or a platform outage.

Because Clef and Clef-Flash handle a 64k-token context, the state you send can be the full conversation, prior related tickets, account details and any attachments, not just a truncated first message.

More context narrows the probabilities in the right direction, and the same pass still returns one answer per question rather than a wall of text to parse.

Thresholds and the review lane

Probabilities let you define a lane for uncertain tickets. If the top category is not clearly ahead, the ticket can go to a human triage queue instead of being auto-routed.

That rule is plain application code. You set a threshold, and you can change it per queue, per customer tier or per time of day without touching the model.

The result is a system that automates the easy majority and hands over the genuinely ambiguous minority with its distribution attached.

You can also tune the lane by time of day. A threshold that is aggressive during business hours, when humans are available to review, can be relaxed overnight so nothing waits unnecessarily.

Consistency across shifts and agents

Human triage drifts. Two agents read the same ticket differently, urgency standards shift between shifts, and the definition of an escalation changes as the team grows.

A decision model applies the same schema and the same thresholds to every ticket, so the standard is explicit and reviewable rather than living in people's heads.

When the policy does change, you change the schema or the threshold, and every future decision follows it immediately.

The audit trail helps here as well. Because every routing decision stores its distribution, you can sample historical tickets and see where the automated standard and the human standard diverged.

Structured extraction for the ticket system

Triage often needs fields, not just labels. The order number, the affected product, the environment and the reproduction steps can all be typed questions in the same request.

That turns a free-text ticket into structured data your systems can act on, which is what makes automation beyond routing possible at all.

Because the answers are constrained by the schema, downstream code can trust their shape before it ever looks at a probability.

Latency at inbox scale

Triage happens on the critical path of the first response. Clef-Flash was built for latency-critical work: published figures put its median latency around 39 ms and its p95 near 122 ms, with input pricing reported close to $0.09 per million tokens.

It runs on about 41 GB of VRAM locally, while Clef needs roughly 85 GB, so a latency-sensitive triage service fits on a single accelerator for teams that must keep customer data inside their own perimeter.

Both models are distributed through Workers AI, Ollama, Hugging Face and OpenRouter, so the same schema can start hosted and move local, or the reverse, without changing the application.

Throughput is usually the lesser constraint. Even at high volume the bottleneck is more often the downstream automation than the scoring call itself.

Try it on your own backlog

The quickest way to see how triage behaves on your tickets is to paste a few into the playground, define the category, urgency and escalation questions, and read the probabilities.

The API documentation describes the request and response shape for wiring it into a real inbox, and the pricing page sets out what a hosted call costs before you commit to anything.

Try it here

Run this through the model

Paste your own state below and define a typed question. The frame is the live Clef-Flash demo; the panel on the playground page calls the API directly.

Clef-Flash community-hosted

Community-hosted demo of the 9B model. Runs live in your browser session.

Open the playground

Open the playground

Read the API docs

View pricing

Prefer code? The API takes the same state and schema. Read the API docs

Frequently asked questions

Can AI triage replace human support agents?

No. Triage decides where a ticket goes and how urgent it is. Humans still resolve it, and probability-based routing keeps the ambiguous cases with people rather than guessing.

How many fields can one triage request return?

As many typed questions as your schema defines. Category, urgency and escalation are typical, and you can add language, product area or an extracted order number in the same pass.

Do I need to train a model on my ticket history?

No. The categories and thresholds are defined in the schema and in your code, so if your policy changes you update the request rather than retraining a model.

Is decision-model triage fast enough for a live inbox?

Clef-Flash is built for latency-critical work, with published median latency near 39 ms and p95 near 122 ms, fast enough to run on every ticket as it arrives.