By late September 2026, builders had a new name for an old job: decision model. TypeSafe's Jev coined the product category. Within weeks the contract showed up as open clones, a local runtime, a Fastino 340M gate, Perplexity's decider, an OpenAI Decisions API preview, and — on October 1, 2026 — Cloudflare Clef and Clef-flash on Workers AI.
This post is the practitioner category guide, not a Kahneman naming essay and not a launch dump. For the System 1 / System 2 metaphor, read What is a System One Model?. For Jev's launch claims, read TypeSafe AI launches Jev. Here the question is architectural: what sits in the hop, what the API actually returns, and when you should keep using LLM classification plus a fitted layer instead.
TL;DR
| Question | Direct answer |
|---|---|
| What is a decision model? | A non-autoregressive classification head. Schema choices are scored from backbone representations in one forward pass. You get calibrated probabilities on typed fields: noul, choice, score. |
| How is that different from JSON mode? | JSON mode still decodes tokens. A decision model never writes the answer as text. It reads a distribution over a closed set. |
| Who named the category? | Diogo Almeida / TypeSafe, via Jev. Clef and Clef-flash speak the same API and ship Apache 2.0 weights. |
| When do I use one? | Closed label set, you need a threshold, and you will call it thousands of times. Agent routing is the canonical fit. |
| When do I skip it? | You need prose, a new option that was not in the schema, or a deterministic audit trail. Also skip if a BERT/XGBoost head already wins on your features — see Jev vs XGBoost and BERT. |
| Are vendor benches enough? | No. Cloudflare's Jev Decision Index has Clef leading several rows. HN testers flagged overfitting, toy loops, and one report (agrippanux) of Clef 2–3× slower and worse on hate-speech than Jev. |
What a decision model actually computes
Start from the inference graph, not the marketing.
A chat LLM is autoregressive: it samples the next token given the tokens so far. Structured-output mode constrains that sampler so the tokens happen to parse as JSON. You still pay decode steps, and the "confidence" you scrape from logprobs is a next-token likelihood, not a calibrated P(class).
A decision model freezes (or lightly adapts) a language-model backbone, then attaches a classification head. The head scores every schema option against the same hidden state. There is no decode loop. The answer is an argmax or a temperature-scaled distribution over labels you enumerated before the call.
That is why the three typed fields showed up everywhere after Jev:
- noul — a boolean framed as a calibrated probability (yes / no, pass / fail, allow / block).
- choice — pick one of N described options. The descriptions matter; near-identical labels behave like a bad softmax.
- score — a numeric value on a declared scale, again with a confidence you are supposed to be able to threshold.
If you only remember one sentence: the model does not write "refund"; it scores the refund class.
The mechanism write-up that matches this picture most closely on explainx.ai is How does Jev work? — parallel scoring over a small output space, not sequential generation. Clef did not invent a second mechanism. It implemented the same hop on a different backbone and host.
The API shape you should design against
Do not design your agent around a vendor SDK name. Design around a contract.
Every production hop we have seen in this category — Jev, OpenJev, Kev, Clef, pplx-decider, Fast-Jev plugins — is some dialect of:
- State — the document, ticket, page, or tool trace you are judging.
- Question — a typed ask (noul / choice / score), not a chat prompt.
- Criteria — optional rubric text the head can attend to.
- Selection — the winning label or number.
- Confidence — a probability you can log, threshold, and override.
- Human override — because the hop is cheap and wrong answers are schema-valid.
A request looks like this in spirit, even when field names drift:
{
"state": { "ticket": "Customer wants a refund on order 1842 after 40 days." },
"questions": [
{
"id": "route",
"type": "choice",
"options": [
{ "id": "refund", "label": "Issue a refund under the 45-day policy" },
{ "id": "deny", "label": "Deny: outside policy or missing evidence" },
{ "id": "escalate", "label": "Escalate to a human reviewer" }
]
},
{
"id": "policy_ok",
"type": "noul",
"prompt": "Does this request fall inside the written refund policy?"
}
]
}
A useful response is not a paragraph. It is route=escalate, p=0.61, plus policy_ok=true at 0.88. Your router code branches on those fields. Your eval harness scores them against a label file. Your on-call playbook says "if confidence is below 0.7, page a human."
That is also why OpenAI's Decisions API is a related product surface and not automatically a decision model. Constraining Luna to emit an enum is still generation unless the serving stack actually replaced the decode loop with a classification head. Read the architecture, not the product name.
Architecture: backbone, LoRA, routing head
The Clef write-up from October 1, 2026 is the clearest public recipe in the category after Jev, so use it as a worked example — not as the only legal architecture.
Clef freezes a Qwen3.8-27B backbone, trains a rank-256 LoRA, and adds a routing / classification head. Context is 64k. A vision encoder is in the stack, so image-bearing state is in-scope in a way early Jev (text-only, 32k in the launch materials) was not. Weights are Apache 2.0. The hosted path is Workers AI. The API is Jev-compatible on purpose: you should be able to retarget a client without rewriting the agent.
Clef-flash uses Qwen3.5-9B on the same product idea. Smaller backbone, same typed fields, meant for the latency-sensitive hop.
Training, as described for Clef, is not "RLHF but faster." The supervised piece is label-smoothed cross-entropy plus a Brier term so the probabilities are not just peaked — they are scored as probabilities. The preference / policy piece is RLCD (Reinforcement Learning for Calibrated Decisions): reward the model when a stated 70% is right about 70% of the time, not when a rater likes the wording (there is no wording) and not when a unit test passes (there is often no unit test).
That objective is why this category keeps colliding with the LLM-as-classifier calibration post. Chat models give you a hard label and a decorative softmax. Decision models claim a Brier-aware head. Claims still need your holdout set.

When to use a decision model vs LLM logprobs
Use this as a working rule, then violate it with data.
Reach for a decision model when:
- The label set is closed and you can write option text a human would understand.
- You will call the hop in a loop — agent routing, tool-call gating, cheap verification checkpoints, compaction, moderation, ticket triage.
- You need a threshold, not a paragraph. Precision/recall is a dial, not a vibe.
- Per-call latency or token cost of a reasoning model is the thing you are actually paying for.
Keep LLM logprobs, or an LLM-plus-logistic wrapper, when:
- You do not have a decision-model vendor or GPU budget yet, but you do have labels. Fit the wrapper. You get calibration without a new serving stack.
- The "options" are not really options — they are open-ended categories you have not enumerated.
- You need the model to explain the call in the same hop. A decision model can score "include rationale" as a separate question; it cannot write the rationale.
- Your existing BERT, GLiClass, or gradient-boosted tree already beats the hosted hop on your features. Jev vs classical classifiers is the honest comparison, not Jev vs GPT on a chatbot bench.
Never use either as a reasoning replacement. Multi-step plans, novel tool invention, and long-horizon coding stay on a System-2 model. The decision model sits beside that loop: "is this the right sub-agent?", "did the last tool call actually finish?", "should we compact this trace?"
Latency, price, and why both vendor tables and HN can be right
Cloudflare's October 1, 2026 table listed Clef-flash at about 38.8ms median versus Jev at about 524ms. The same post's threat-intel example timed domain classification at 2.2s on Clef versus 4.7s on gpt-oss-120b. Those are vendor numbers on vendor workloads. They tell you the shape of the win — classifier hop vs generating a label with a 120B decoder — not what your p95 will be on Workers AI from your region with a 8k-token state.
Hacker News users reported the opposite on hosted Clef: slower than Jev in real tests. agrippanux's write-up in that thread is the one to remember: Clef 2–3× slower than Jev in their setup, and worse on a hate-speech task. Both statements can be true at once. A median on a warm, short-prompt microbench is not a hate-speech eval with long comments and a cold start.
Pricing, as cited on HN around that discussion:
| Hosted hop | Input price (per million tokens) | 1M decisions at 300 tokens / call |
|---|---|---|
| Jev | ~$0.042 | ~$12.60 |
| Clef-flash | ~$0.09 | ~$27 |
| Clef | ~$0.24 | ~$72 |
Output tokens are typically free in this category because there are no output tokens in the generative sense. That is the whole point. Still budget the input: a 64k context is a feature until someone pastes a PDF into state on every call.
If you are choosing a host, log three numbers for a week on your distribution: p50, p95, and dollars per thousand decisions. Ignore "200× cheaper than an LLM" until you name the LLM, the prompt, and the decode length.
Benchmarks, clones, and the skepticism that is now part of the category
Cloudflare published a Jev Decision Index on which Clef leads several rows. Treat that the way we treated JevBench: useful as a vendor panel, insufficient as a purchase order. The HN skepticism is specific, not vibes:
- Public benches leak. A LoRA on a frozen Qwen can look brilliant on a set the community already overfit.
- Toy loops (the Pokemon-style examples that circulated with the Clef discussion) measure "can the head pick a legal action," not "will this route tickets for six months."
- Independent testers can — and did — reverse the latency and quality ranking on a task that was not in the index.
The clone wave is the other half of the category, and it is why a durable guide cannot be "just use Jev":
- OpenJev — browser-side readout so you can feel one-pass scoring without a vendor key.
- Kev — the most documented Qwen3.5-family open counterpart.
- Laya — on-device MLX; viral speed claims, check the project's own benches.
- Julia-1 — CPU-first, useful when the hop has to live next to a request thread.
- GLiNER 2.5 Decide — 340M open weights for cheap gating.
- Perplexity pplx-decider — Apache-2.0 27B-class weights plus a $0.04/M Decisions API.
- Fast-Jev — a plugin that uses noul decisions to prune traces instead of summarizing them.
- Jev Ultrafast / browser-use — typed decisions inside a browser agent, the "routing as UI" version of the same hop.
The six clones in two days post is still the best map of how fast the API ossified. Once the contract is stable, the differentiator is calibration on your labels, not who coined "System One."
Agent routing: the job that justifies the category
If you only adopt one pattern, adopt this one.
An agent loop is mostly small, repeated judgments: which tool, which sub-agent, whether the last observation is enough, whether to stop. Those judgments do not need a 200-token chain of thought. They need a closed set and a number you can threshold.
Wire it like this:
- Keep the writer / planner on a reasoning model.
- Put a decision-model client behind one interface (
decide(state, questions)). - Start with noul gates ("should we call the browser?") before you try a 20-way choice.
- Log every selection and confidence. Sample overrides weekly. That log is your fine-tune set.
- Fail open or fail closed on purpose. A 0.51 noul is not a policy.
Worked integration notes live in How to integrate Jev for agent routing. Swap the base URL later — Clef, pplx-decider, Kev, a local Ollaya process — without rewriting the planner.
Honest limits
Not deterministic. Same state, same schema, two calls, two answers happens. If you need a hash, use a rule.
Schema-valid ≠ true. "Never hallucinates" in launch copy means "cannot emit an illegal enum." It does not mean the refund was correct.
Option text is feature engineering. Choice heads discriminate described options. Sixteen well-written actions beat ninety near-identical grid cells. That failure mode showed up in independent Battleship tests around the Jev launch, and it will show up in your router if you dump raw IDs into label.
Vision and context are product features, not category laws. Clef's 64k and vision encoder matter if your state is a screenshot. They do not make every clone multimodal.
Calibration is an eval, not a flag. Brier loss in training is necessary, not sufficient. Plot reliability on your holdout. If 0.9 is right 70% of the time, your threshold is a lie.
Hosted latency is a distribution. Cloudflare's 38.8ms and HN's "Clef felt slower than Jev" can coexist. Measure.
What to do on Monday
- List the hops in your agent that are already enums in disguise.
- Write the schema first. If you cannot name the options, you do not have a decision-model task.
- Prototype with whatever Jev-compatible endpoint you already have. Log confidence.
- Compute Brier and a confusion matrix on 200 real examples before you switch production traffic.
- Keep the LLM for the hops that write. Put the decision model on the hops that choose.
Related reading
- What is a System One Model? — Kahneman naming, complementary to this architecture guide
- TypeSafe AI launches Jev — the product that named the category
- How does Jev work? RLCD explained — parallel pass vs decode
- Stop using LLMs as classifiers — logistic wrapper when you are not ready to switch models
- How to integrate Jev for agent routing — the production hop
- Six Jev clones in two days — how the API spread
- Perplexity pplx-decider and Decisions API — open weights at $0.04/M
- Jev vs XGBoost and BERT — when classical ML still wins
- Earendil π 1: a durable minimal harness · DeepSeek Harness desktop — where typed hops sit next to a local agent loop
Architecture, API shape, and the Clef / Clef-flash figures above reflect TypeSafe's September 2026 category launch plus Cloudflare's October 1, 2026 Clef materials and the Hacker News discussion that followed, as of October 2, 2026. Hosted latency, index rows, and list prices move. Verify on your own traffic before you rip out an LLM hop.
