Why a binary verdict is not enough
Most moderation endpoints answer one question: is this allowed or not. That looks simple until you have to explain a decision to a user or tune it for a specific community.
A single verdict hides the cases that matter most. Borderline content, satire, quoted speech and context-dependent posts all get flattened into the same yes or no, and you have no handle to adjust without changing the provider.
An AI moderation API built on a decision model returns the whole picture instead: a probability for each label you defined, in one pass, so the threshold lives in your code and not in someone else's black box.
Appeals make the point concrete. When a creator disputes a takedown, a probability distribution lets you show which signals were weak and by how much, which is far more defensible than pointing at an opaque label.
The shape of a moderation decision
A decision model reads a state — the post, comment, caption or image — together with a schema of typed questions, and returns a probability for every allowed option. It does not write a ruling in prose.
Clef and Clef-Flash are the open-source decision models behind this approach, released by Cloudflare under Apache 2.0. Clef is the 27B precision model; Clef-Flash is the 9B latency model. Both handle a 64k-token context and include a vision encoder.
For moderation the schema is usually an enumeration of actions such as allow, review and block, plus any additional typed questions you want answered at the same time.
Allow, review and block as probabilities
Instead of one verdict you receive three numbers. A score of 0.88 for review, 0.09 for allow and 0.03 for block tells a very different story from a post that scores 0.55, 0.40 and 0.05.
That detail lets you build a graded response. Content that is clearly fine ships immediately, content that is clearly harmful is blocked, and the ambiguous middle goes to a human queue with its distribution attached.
Because the probabilities come from one evaluation of the same state, they are comparable with each other, which a set of independent prose answers would never be.
Thresholds are policy, not model
The useful consequence is that the line between block and review is yours to draw. A children's platform can set a stricter cutoff than a developer forum, using the same model and the same schema.
Changing that policy is a configuration change rather than a retraining job. When a new kind of abuse appears, you adjust the threshold or add a category, and you can ship the update the same day.
You can also keep separate thresholds per surface, so comments, direct messages and public posts each get the sensitivity they deserve without running different models.
This is also what makes the approach testable. You can replay last month's content against a new threshold and see exactly how many items would move from allow to review, or from review to block, before any policy goes live.
Several labels in one pass
Moderation is rarely one question. You may need a category, a severity and a confidence-independent action, all for the same item.
A decision model answers all of them from a single reading of the state. You send one schema with an enumeration for category, a bounded number for severity and an enumeration for the recommended action.
The answers arrive together and stay consistent, because they were scored from one interpretation rather than three independent calls that can disagree with each other.
Consistency also removes a subtle failure mode. When category and action are decided by different calls, a model can label something clearly harmful and then recommend allowing it, because the second call never saw the first answer.
Severity as a bounded number
Not every decision is a label. Severity is naturally a number, and a typed question can bound it to a range so the model cannot return a value outside it.
A severity score lets you sort the human review queue, so moderators see the most harmful items first instead of an unordered pile.
It also lets you express escalation policy in plain arithmetic: block above one cutoff, queue above a lower one, and let the rest through.
Text, JSON and images together
Modern moderation has to see more than text. A post is often an image, a caption and structured metadata at the same time.
Clef and Clef-Flash include a vision encoder in the base model rather than as a bolt-on, so the same request can carry a screenshot, a document and a JSON record without a separate pipeline.
That single interface keeps the moderation stack small, which matters when you have to reason about failure modes and audit how a decision was reached.
Latency and cost at moderation scale
Moderation runs on every piece of user content, so latency and price decide whether an approach is viable. Clef-Flash was built for latency-critical work: published figures put its median latency around 39 ms and its p95 near 122 ms, with input pricing reported close to $0.09 per million tokens.
It runs on about 41 GB of VRAM locally, while Clef needs roughly 85 GB, so the latency model fits comfortably on one accelerator for teams that must keep user content inside their own perimeter.
Both models are distributed through Workers AI, Ollama, Hugging Face and OpenRouter, so a team can start hosted and move to local deployment, or the reverse, without changing the schema.
Building the human review queue
Probability output turns a review queue into something you can prioritise rather than merely populate. Sort by the block score, group by category and show the distribution next to each item.
Reviewers can then confirm or overturn decisions, and those outcomes let you measure whether your thresholds are set where you think they are.
Because the routing rules live in your code, tightening or loosening them after a review cycle is a normal pull request rather than a conversation with a vendor.
Over time the queue becomes a labelled dataset as a side effect, and you can use it to check whether a different threshold would have served your community better.
Start scoring your content
The fastest way to see how this behaves on your own material is to paste a sample into the playground, define allow, review and block, and read the probabilities directly.
The classification API page describes the request and response in detail, and the pricing page sets out what a hosted call costs before you integrate anything.