explainx / blog / topics
Decision Models
Decision models are small models trained to make one call well: approve or reject, route or escalate, pick an option. They are cheaper and faster than general chat models for that job.
This page tracks the category as it forms: new models and APIs, benchmark results, open-weight releases, and how teams use them in production.
34 stories · latest Oct 8, 2026
Start here
Top 10 Decision Models in October 2026: Jev, Perplexity, OpenAI and Open Source
In three weeks the decision-model category went from one hosted product to a crowded field of hosted APIs and open weights. This is a ranked list of ten you can use today, starting with Jev, with prices, latency, licenses and the caveat that matters most for each.
What Are Decision Models? The Practitioner Category Guide
A decision model is not a chat LLM with JSON mode. It is a non-autoregressive classification head that scores a schema of choices from backbone representations in one forward pass and returns calibrated typed fields. This guide covers architecture, API shape, calibration, latency, clones, and when to keep using LLM logprobs instead.
How GLM-5.3-Flash Scores a Closed Choice in One Token
If you already serve a generative model and the product mostly asks closed questions, you can score the option list from one token's log probabilities instead of paying for a JSON decode. Edgeless Systems did that with GLM-5.3-Flash and published a head-to-head against Jev and Laya.
A Jev-like LLM wrapper using logprobs — including vision — explained
Instead of a dedicated decision model, you can ask a chat model for a single letter answer with logprobs enabled and treat relative token probabilities as calibrated-ish scores. Allan Boll's September 2026 blog post and HN thread show a ~100-line Python pattern that works on llama.cpp and OpenAI — including Gemma 4 vision on an RTX 3090 at about 1 FPS for three questions per frame.
LLMs Repeat 96% of Jev-Style Confident Errors, Undermining Eval Cascades
A cheap classifier makes a confident, wrong call. An LLM is supposed to catch it on the second pass. A research finding making the rounds puts a number on how often that second pass actually works — and it's not good: 96% of the time, the LLM agrees with the classifier's confident mistake instead of correcting it.
Ollaya Is "Ollama for Decision Models" — A Local Runtime, Not a New Model
Every Jev clone so far has shipped a single model. Ollaya ships none of its own — instead it's a desktop app, CLI, and Docker image that bundles seven open decision models behind a drop-in TypeSafe-compatible local endpoint, the same category move Ollama made for local LLMs.
Timeline
October 2026
Oct 8
Top 10 Decision Models in October 2026: Jev, Perplexity, OpenAI and Open SourceOct 7
Strands Decider 2B: A Small Open Decision Model You Can Run LocallyOn October 1, 2026 the Strands team published Strands Decider 2B, a 1.9B open decision model that picks between options and returns a confidence with every answer. Code and weights are Apache-2.0, and the training data sources and scripts are public. Here is how the pointer head works, what the benchmarks do and do not say, how it stacks up against Jev, pplx-decider and OpenAI Decisions, and how to run it today.
Oct 6
Liquid AI d1 Now Sees Images: A $0.04 Decision Model vs GPT-6.1 SolLiquid AI's d1 is a decision model: it answers yes/no, multiple-choice and score questions by returning calibrated probabilities instead of generating text. The new version accepts images, and Liquid says it matches or beats GPT-6.1 Sol on four of six real tasks at 19 to 200 times lower cost. Here is what that does and does not mean.
Oct 2
What Are Decision Models? The Practitioner Category Guide
September 2026
Sep 29
GPT Researcher 3.7.0: Jev Scores Passages by UsefulnessAssaf Elovic's GPT Researcher tagged v3.7.0 on September 26, 2026 with a Jev context filter and a BM25 default when you have no TypeSafe key. Aggregators flattened a 73% kept-passage precision number into a quality headline; the primary eval is a 28-task replay, and embeddings are still optional rather than replaced.
Sep 29
Jeff: Home-Trained Jev-Compatible 0.8B Decision ModelsA Hacker News front-page project from firelex ships Qwen3.5 and Gemma fine-tunes that speak Jev's request format locally. The 2B's five-benchmark panel score is 83.1% against Jev's published 83.0% on a different sample; the 0.8B is ~28 ms on an M4 Max. This post is about Jeff specifically — not Kev, OpenJev, or Laya — and what the thread actually argued.
Sep 27
How GLM-5.3-Flash Scores a Closed Choice in One TokenSep 27
Julia 1: A 144M Decision Model You Can Run on a CPUSupersonic Labs released Julia 1, a 144.3 million parameter encoder you can run on a CPU or in a WebGPU browser under Apache 2.0. The published edge over a Jev reference is 0.45 points on one suite, and Banking77 falls to 64%.
Sep 26
A Jev-like LLM wrapper using logprobs — including vision — explainedSep 26
LLMs Repeat 96% of Jev-Style Confident Errors, Undermining Eval CascadesSep 26
Ollaya Is "Ollama for Decision Models" — A Local Runtime, Not a New ModelSep 26
Respan Span-01: hyper-parallel behavior classifier vs Jev — what the launch claimsRespan AI's Span-01 promises frontier-grade behavior detection — prompt injection, tool misuse, secrets, agent loops — at classifier speed and roughly half Jev's price with an 18% benchmark lift, plus a free Lite tier. Y Combinator amplified the thread. explainx.ai maps the architecture claims, how they differ from Jev's sampler, and what to verify before swapping verification checkpoints in production agents.
Sep 22
LangSmith Adds Jev to Score Production Agent TracesLangChain shipped Jev-as-a-judge inside LangSmith Evals: attach typed Choice, Score, and Noul questions to production traces, run online evaluators on live traffic, and route calls through LangSmith Gateway with guardrails and cost accounting. Here's how it differs from the offline benchmark and what to configure first.
Sep 22
SemIf on LangSmith Gateway: Free Decision AI Through Sept 28LangChain added a Decision models category to the LangSmith LLM Gateway on September 22, 2026, with hosted SemIf (semif-qwen3.5-4b) free through September 28 on US Free, Developer, and Plus workspaces. Here's what SemIf is, how its authored144 benchmark fits the Jev ecosystem, and what you get versus bringing your own TypeSafe (Jev) key.
Sep 21
DocJev: LlamaIndex's Jerry Liu Puts Jev on Document Classification and SplittingJerry Liu, LlamaIndex's cofounder and CEO, released DocJev — an open-source library that hands document classification and document-splitting decisions to Jev instead of a general-purpose LLM. The published benchmark shows classification dropping from 794ms to 138.6ms median latency, but a replier's qualifier about the accuracy pilot's small sample size is worth reading before trusting it unattended.
Sep 21
Using Jev as Cheap Verification Checkpoints in Agent PipelinesA checkpoint that costs a fraction of a cent only pays for itself if it changes what happens next. This guide works through where to place Jev checks in a research-to-article agent pipeline, the real cost math behind "cheap enough to check constantly," and the honest failure modes — noisy alarms, distracting context, and checks with no attached action — that make a checkpoint worthless even when it's nearly free.
Sep 21
Jev Is Now Open to Everyone — No WaitlistSix days after launching to a 140,000-signup waitlist, TypeSafe AI opened Jev to everyone on September 21, 2026 — no approval, no queue. New accounts start with $5 in credit, which TypeSafe says is worth roughly 120 million tokens. Here's what changed, what that credit is actually worth against Jev's own disclosed accuracy numbers, and what to check before you build on it.
Sep 21
Jev Playground and JevBench: What TypeSafe AI Actually ClaimedTwo TypeSafe AI headlines hit an AI news digest on September 20-21, 2026: a Jev Playground claiming "440x cheaper than LLMs," and a new decision-model benchmark, JevBench, with Jev reportedly leading at 75.3. Neither claim comes with a linked source article, and the 440x figure is already the second different "Nx cheaper" number TypeSafe has published in a week.
Sep 21
Jev Ultrafast: Browser Use Puts Jev in the Browser Agent LoopBrowser Use, the team behind the popular browser-use agent library, shipped Jev Ultrafast — an open-source browser agent that reads a structured element table instead of screenshots and lets Jev pick an operation and a target element per step, with a small LLM only invoked to write text. The published demo completes a real Google Flights search in 7.1 seconds, with independent outcome verification.
Sep 21
Kev's Real Numbers: Inside the Open-Source Jev Clone's 0.8B/4B/9B Familyexplainx.ai previously covered Kev only through unverified digest headlines — a 0.5B MacBook model, then an "8B" follow-up with no source. The project's actual GitHub README and Hacker News launch thread are now public, and they describe something different: a documented 0.8B/4B/9B family with LoRA adapters, a pointer-head architecture, and benchmark numbers run against Jev on both trained and unseen data.
Sep 21
Laya-MLX: A Real On-Device Alternative to Jev — Is It Really 50x Faster?Developer mizorewww's laya-mlx is a real, open-source Apache-2.0 MLX port of Convai Innovations' Laya typed-decision model, with published benchmarks showing sub-14ms decisions and under 1GB peak memory on an M3 Max. A viral Chinese-language X post calls it "50x faster than Jev" — a claim the project's own README never makes. Here's what's actually measured, what isn't, and how it fits next to TypeSafe AI's Jev.
Sep 21
TypeSafe's Founder Published Coding-Agent Notes. The KV-Cache Math Is the Part Worth Reading.TypeSafe AI founder Diogo Almeida published a long, explicitly speculative notes document on what a Jev-centric coding agent could look like — and hopes the community builds it before he does. The most concrete, checkable claim inside is a worked cost comparison showing that routing a task to a cheaper model and back to a stronger one can cost more than never switching, because the stronger model has to reprocess the whole context from scratch.
Sep 20
Awesome Jev Use Cases: A 50-Demo Gallery You Can Run YourselfEvery Jev use-case argument so far has been reasoning about the shape of the Choice, Score, and Noul primitives. The awesome-jev-use-cases repo skips the reasoning and ships 50 runnable demos instead — each one a side-by-side comparison against OpenAI's Responses API with a live 2D visualization, no API key needed until you want your own numbers.
Sep 20
Bespoke Nimble: A 9B Model Hit 90% on Jev — Built in Days, Not MonthsA small, independently-built open model reportedly closed most of the gap to TypeSafe AI's Jev on its own eval, in a fraction of the time a frontier lab spends on a launch. The real story isn't the score — it's what a 9B LoRA fine-tune beating expectations says about efficient fine-tuning versus brute-force scale, and how skeptically to read any "hits X% on eval Y" headline.
Sep 20
Jev vs LLM-as-Judge: LangChain Benchmarks Agent EvaluationLangChain ran the same Deep Agents weather-tool traces through four judges — TypeSafe AI's Jev, GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 — and measured accuracy against a human oracle, per-case variance, cost, and latency. Jev matched the human oracle on all 500 repeated decisions at roughly 1/80,000th the cost of Claude.
Sep 20
Top 10 Ways to Learn Jev (TypeSafe AI) in 2026: Courses, Workshops, and ResourcesJev, TypeSafe AI's "System One Model," is five days old and search results for "Jev course" are already a mess of speculation. Here's an honest, ranked list of the real resources worth your time — starting with explainx.ai's own self-paced Udemy course and live Jev workshop.
Sep 19
How to Wire Jev Into Your Agent Pipeline for Routing DecisionsJev is available directly on Vercel's AI Gateway, exposed through AI SDK 7's experimental_evaluate function, and has an official LangChain integration (TypeSafeClassifier) built specifically for routing, escalation, and tool-call decisions inside an agent loop. Here's how to actually wire it in, with the concrete integration points and what each one is for.
Sep 19
Jev's Actual Security Use Case: Detecting Prompt Injection, Not Getting HackedThere's no published adversarial research on gaming or poisoning Jev, TypeSafe AI's non-generative "System One Model" — a search for that angle comes up thin. What does exist is the inverse: Jev being positioned as a security tool itself, with a `contains_prompt_injection` classification primitive meant to sit in front of a main LLM and flag jailbreak or injection attempts fast and cheap, before they reach the model actually generating your response.
Sep 19
Is Jev's 200x-Faster, 400x-Cheaper Claim Actually True?TypeSafe AI's headline numbers for Jev — 20-200x faster, 40-400x cheaper than LLMs on structured-output tasks — are TypeSafe's own benchmarks, measured against agreement with other frontier models rather than verified ground truth. An independent test from Every corroborated the general direction but called results "good but not perfect," and Jev's own dashboard shows a real accuracy gap against the best comparator model.
Sep 19
Jev vs. XGBoost and BERT: Is a System One Model Actually New?Before Jev, teams needing fast structured classification typically reached for XGBoost (fast, cheap, lower ceiling on accuracy) or a fine-tuned BERT model (higher accuracy, more setup, still not free-text generation). Jev sits in a genuinely different spot on that spectrum — not because typed-output classification is new, but because of how it's trained and how it reports confidence. Here's an honest comparison.
Sep 19
OpenJev: Try Jev's Trick Yourself, Free, in Your BrowserA developer built OpenJev — a free, entirely in-browser tool that lets anyone run open models like Qwen3 and MiniCPM5 locally and directly compare TypeSafe's Jev-style "direct readout" decision method against ordinary token-by-token generation, on their own GPU. It hit 556 points on Hacker News, and the discussion is as much about a naming dispute and a vibecoded-looking UI as it is about the underlying technique.
Sep 19
Six Jev Clones Shipped in Two Days — Here's What Each One Actually DoesTypeSafe AI's Jev launched September 15, 2026. Within two days, at least six independent open-source clones or alternatives appeared, catalogued by Latent.Space — ranging from a 421M-parameter ModernBERT-based model to a 40KB embedding-only implementation to a 0.5B model designed to run on a MacBook Pro. Here's what each one actually is, and what the speed of the response says about how replicable Jev's core idea turned out to be.
Sep 19
Where Jev Actually Fails: The Specific Complaints Behind the HypeJev's Hacker News launch thread ran to 256 comments, and buried in the general skepticism are specific, concrete failure modes worth taking seriously — not "it's not an LLM" complaints, but named cases where Jev returns a type-valid, well-formed, confidently-scored answer that is simply wrong. Here's what's actually been reported, sourced directly.
Sep 17
Could Jev Run Self-Driving? The Latency Argument, and Why Engineers Pushed Back"The implications of Jev on self-driving could be huge" drew 209,000 views and an immediate wall of pushback from people who work on autonomy. Their objections are specific and they are right, but the underlying question of where a fast decision model belongs in a robotics stack is still a good one.