Decision Models

Decision Models are AI models that return a typed decision, such as one label from a fixed list, a score or a yes/no, together with a calibrated probability, instead of generating free-form text. They emerged as a product category in late September and early October 2026, when TypeSafe opened early access to Jev, OpenAI previewed a Decisions API and Cloudflare open-sourced Clef, all within a few days of each other.

Why a Separate Category

A large share of what software asks a large language model to do is not writing at all. It is choosing: which tool to call, which queue a support ticket belongs in, whether a request is safe, whether a crawler is a good bot or a bad one. With a general LLM, developers get those choices by prompting for text, parsing the output and hoping it matches the expected format. Every call pays for token-by-token generation, and the model's stated confidence is usually just more text. Decision models flip this around. The developer supplies the question and the set of allowed answers, and the model returns one of those answers along with a probability that is meant to be trustworthy enough for software to act on directly. Cloudflare's definition is simple: "A decision model makes classifications to help agents decide how to act, based on certain probabilities."

How They Work

Scoring instead of generating

An ordinary LLM produces output autoregressively, one token at a time. A decision model instead scores the fixed set of valid answers in parallel. Cloudflare describes Clef's decision step as "non-autoregressive, so there's no intermediate text to generate token by token"; the model "scores the valid schema choices in parallel." Because the answer can only ever be one of the allowed values, the output always matches the schema. TypeSafe says Jev "never makes type errors", and its own fine print notes that its "0%" type-error figure is "not empirical. Schema matching is guaranteed." This is also why the vendors can claim large speed gains: there is little or no output to generate.

Calibration

The second ingredient is calibration, meaning that when the model says 80%, it should be right about 80% of the time. TypeSafe describes its stack as "a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD)." Cloudflare's Clef post-training uses "label-smoothed cross-entropy for valid schema outputs paired with a Brier loss to refine probability calibration," plus RLCD. The Brier score penalizes the squared gap between a predicted probability and the actual outcome, so training against it pushes stated confidence toward observed accuracy. Calibrated probabilities let developers set thresholds, for example acting automatically above 95% and escalating to a human or a larger model below it.

The Three Implementations

TypeSafe Jev

TypeSafe, founded by Diogo Almeida, introduced "System One Models", which it calls "a new class of frontier models built to make fast, structured decisions that software can use directly." Jev is the first public one, and early access opened on September 28, 2026. TypeSafe claims end-to-end response times of "70ms-500ms", which it says can be "40x-200x faster" than frontier LLMs, while cautioning that it expects these to be "on the higher end of real world gains." It prices input at $0.042 per million tokens and makes output free. Its docs list three primitives: Choice, Score and Noul.

OpenAI Decisions API

At DevDay on September 29, 2026, OpenAI reportedly put a Decisions API into limited preview, built on GPT-6 Luna. The developer supplies a question and the allowed answers. According to a Firecrawl comparison, OpenAI claims about 150 ms per decision versus about 1.6 seconds for a standard call. No schema or pricing has been published.

Cloudflare Clef and Clef-flash

On October 1, 2026, Cloudflare released Clef and Clef-flash as open-source decision models under the Apache 2.0 license. Clef is built on a Qwen3.8-27B backbone and Clef-flash on Qwen3.5-9B. They add a vision encoder and a 64k context window (Jev's is 32k). Cloudflare reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash, against 524.1 ms for Jev in its own testing. Both are generally available on Workers AI.

Jev's API as a De Facto Interface

The most telling detail in Cloudflare's launch is that Clef is "fully Jev-API compatible, so you can experiment with these hosted models easily." A request format a startup published only days earlier has become the interface a major cloud provider chose to match, much as OpenAI's chat completions format became the default for LLMs. A community benchmark, the Jev Decision Index, has grown up around the same interface, and Cloudflare says "Clef is currently the leader" on it. That is Cloudflare's claim about a community leaderboard, not an independent audit.

The "System One" Framing

"System One" is TypeSafe's term. It is "inspired by Daniel Kahneman, Thinking, Fast and Slow," drawing on "the distinction between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning." In that framing, reasoning models and chain-of-thought are System 2: slow, expensive and deliberate. Decision models are System 1: quick, cheap judgments made many times. In practice an agentic system might use both, with a decision model triaging and routing, and a frontier model handling the hard cases.

Use Cases

  • Routing. Picking which model, tool or agent should handle a request, a natural fit for agent harnesses and AI gateways.
  • Classification. TypeSafe pitches Jev for tasks that "classify, route, score, extract, or branch"; Cloudflare uses Clef internally for domain classification and support triage.
  • Guardrails. TypeSafe also pitches it to "score, judge, verify, guardrail, and detect jailbreaks", relevant to prompt injection defenses and content moderation. Cloudflare uses Clef for trust and safety review.
  • Bot detection. Clef is "built-in to our Bot products to decide if a crawler is a good bot or bad bot," according to Cloudflare.
  • Map-reduce over large datasets. TypeSafe lists "map-reducing over big data" as a use case: at sub-cent prices and sub-second latency, it becomes practical to run a decision over every row of a large corpus, such as tagging millions of documents, and then aggregate the results.

Limits and Open Questions

Nearly every performance figure in this category comes from a vendor. DataCamp's assessment noted that "no large-scale independent reproduction has surfaced yet" and that the benchmarks are vendor-reported. It also pointed out that TypeSafe "can't prove the pricing isn't subsidized," which leaves open whether free output and $0.042-per-million input are sustainable. innFactory argued that the underlying classifier principle "is not new"; what's new is packaging it as a general-purpose, schema-driven API on top of a large pretrained model. Cloudflare's head-to-head latency numbers compare its own hosted models with a competitor's API, so network placement matters. And OpenAI's offering has no published schema or price yet. On Hacker News the reaction has mostly been people building their own copies, which shows interest but isn't evidence of quality. The next few months of independent testing will show whether calibrated decisions hold up outside vendor benchmarks.

Further Reading