explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

← All topics

explainx / blog / topics

Decision Models

Decision models are small models trained to make one call well: approve or reject, route or escalate, pick an option. They are cheaper and faster than general chat models for that job.

This page tracks the category as it forms: new models and APIs, benchmark results, open-weight releases, and how teams use them in production.

34 stories · latest Oct 8, 2026

Start here

Top 10 Decision Models in October 2026: Jev, Perplexity, OpenAI and Open Source

In three weeks the decision-model category went from one hosted product to a crowded field of hosted APIs and open weights. This is a ranked list of ten you can use today, starting with Jev, with prices, latency, licenses and the caveat that matters most for each.

What Are Decision Models? The Practitioner Category Guide

A decision model is not a chat LLM with JSON mode. It is a non-autoregressive classification head that scores a schema of choices from backbone representations in one forward pass and returns calibrated typed fields. This guide covers architecture, API shape, calibration, latency, clones, and when to keep using LLM logprobs instead.

How GLM-5.3-Flash Scores a Closed Choice in One Token

If you already serve a generative model and the product mostly asks closed questions, you can score the option list from one token's log probabilities instead of paying for a JSON decode. Edgeless Systems did that with GLM-5.3-Flash and published a head-to-head against Jev and Laya.

A Jev-like LLM wrapper using logprobs — including vision — explained

Instead of a dedicated decision model, you can ask a chat model for a single letter answer with logprobs enabled and treat relative token probabilities as calibrated-ish scores. Allan Boll's September 2026 blog post and HN thread show a ~100-line Python pattern that works on llama.cpp and OpenAI — including Gemma 4 vision on an RTX 3090 at about 1 FPS for three questions per frame.

LLMs Repeat 96% of Jev-Style Confident Errors, Undermining Eval Cascades

A cheap classifier makes a confident, wrong call. An LLM is supposed to catch it on the second pass. A research finding making the rounds puts a number on how often that second pass actually works — and it's not good: 96% of the time, the LLM agrees with the classifier's confident mistake instead of correcting it.

Ollaya Is "Ollama for Decision Models" — A Local Runtime, Not a New Model

Every Jev clone so far has shipped a single model. Ollaya ships none of its own — instead it's a desktop app, CLI, and Docker image that bundles seven open decision models behind a drop-in TypeSafe-compatible local endpoint, the same category move Ollama made for local LLMs.

Timeline

October 2026

  1. Oct 8

    Top 10 Decision Models in October 2026: Jev, Perplexity, OpenAI and Open Source
  2. Oct 7

    Strands Decider 2B: A Small Open Decision Model You Can Run Locally

    On October 1, 2026 the Strands team published Strands Decider 2B, a 1.9B open decision model that picks between options and returns a confidence with every answer. Code and weights are Apache-2.0, and the training data sources and scripts are public. Here is how the pointer head works, what the benchmarks do and do not say, how it stacks up against Jev, pplx-decider and OpenAI Decisions, and how to run it today.

  3. Oct 6

    Liquid AI d1 Now Sees Images: A $0.04 Decision Model vs GPT-6.1 Sol

    Liquid AI's d1 is a decision model: it answers yes/no, multiple-choice and score questions by returning calibrated probabilities instead of generating text. The new version accepts images, and Liquid says it matches or beats GPT-6.1 Sol on four of six real tasks at 19 to 200 times lower cost. Here is what that does and does not mean.

  4. Oct 2

    What Are Decision Models? The Practitioner Category Guide

September 2026

  1. Sep 29

    GPT Researcher 3.7.0: Jev Scores Passages by Usefulness

    Assaf Elovic's GPT Researcher tagged v3.7.0 on September 26, 2026 with a Jev context filter and a BM25 default when you have no TypeSafe key. Aggregators flattened a 73% kept-passage precision number into a quality headline; the primary eval is a 28-task replay, and embeddings are still optional rather than replaced.

  2. Sep 29

    Jeff: Home-Trained Jev-Compatible 0.8B Decision Models

    A Hacker News front-page project from firelex ships Qwen3.5 and Gemma fine-tunes that speak Jev's request format locally. The 2B's five-benchmark panel score is 83.1% against Jev's published 83.0% on a different sample; the 0.8B is ~28 ms on an M4 Max. This post is about Jeff specifically — not Kev, OpenJev, or Laya — and what the thread actually argued.

  3. Sep 27

    How GLM-5.3-Flash Scores a Closed Choice in One Token
  4. Sep 27

    Julia 1: A 144M Decision Model You Can Run on a CPU

    Supersonic Labs released Julia 1, a 144.3 million parameter encoder you can run on a CPU or in a WebGPU browser under Apache 2.0. The published edge over a Jev reference is 0.45 points on one suite, and Banking77 falls to 64%.

  5. Sep 26

    A Jev-like LLM wrapper using logprobs — including vision — explained
  6. Sep 26

    LLMs Repeat 96% of Jev-Style Confident Errors, Undermining Eval Cascades
  7. Sep 26

    Ollaya Is "Ollama for Decision Models" — A Local Runtime, Not a New Model
  8. Sep 26

    Respan Span-01: hyper-parallel behavior classifier vs Jev — what the launch claims

    Respan AI's Span-01 promises frontier-grade behavior detection — prompt injection, tool misuse, secrets, agent loops — at classifier speed and roughly half Jev's price with an 18% benchmark lift, plus a free Lite tier. Y Combinator amplified the thread. explainx.ai maps the architecture claims, how they differ from Jev's sampler, and what to verify before swapping verification checkpoints in production agents.

  9. Sep 22

    LangSmith Adds Jev to Score Production Agent Traces

    LangChain shipped Jev-as-a-judge inside LangSmith Evals: attach typed Choice, Score, and Noul questions to production traces, run online evaluators on live traffic, and route calls through LangSmith Gateway with guardrails and cost accounting. Here's how it differs from the offline benchmark and what to configure first.

  10. Sep 22

    SemIf on LangSmith Gateway: Free Decision AI Through Sept 28

    LangChain added a Decision models category to the LangSmith LLM Gateway on September 22, 2026, with hosted SemIf (semif-qwen3.5-4b) free through September 28 on US Free, Developer, and Plus workspaces. Here's what SemIf is, how its authored144 benchmark fits the Jev ecosystem, and what you get versus bringing your own TypeSafe (Jev) key.

  11. Sep 21

    DocJev: LlamaIndex's Jerry Liu Puts Jev on Document Classification and Splitting

    Jerry Liu, LlamaIndex's cofounder and CEO, released DocJev — an open-source library that hands document classification and document-splitting decisions to Jev instead of a general-purpose LLM. The published benchmark shows classification dropping from 794ms to 138.6ms median latency, but a replier's qualifier about the accuracy pilot's small sample size is worth reading before trusting it unattended.

  12. Sep 21

    Using Jev as Cheap Verification Checkpoints in Agent Pipelines

    A checkpoint that costs a fraction of a cent only pays for itself if it changes what happens next. This guide works through where to place Jev checks in a research-to-article agent pipeline, the real cost math behind "cheap enough to check constantly," and the honest failure modes — noisy alarms, distracting context, and checks with no attached action — that make a checkpoint worthless even when it's nearly free.

  13. Sep 21

    Jev Is Now Open to Everyone — No Waitlist

    Six days after launching to a 140,000-signup waitlist, TypeSafe AI opened Jev to everyone on September 21, 2026 — no approval, no queue. New accounts start with $5 in credit, which TypeSafe says is worth roughly 120 million tokens. Here's what changed, what that credit is actually worth against Jev's own disclosed accuracy numbers, and what to check before you build on it.

  14. Sep 21

    Jev Playground and JevBench: What TypeSafe AI Actually Claimed

    Two TypeSafe AI headlines hit an AI news digest on September 20-21, 2026: a Jev Playground claiming "440x cheaper than LLMs," and a new decision-model benchmark, JevBench, with Jev reportedly leading at 75.3. Neither claim comes with a linked source article, and the 440x figure is already the second different "Nx cheaper" number TypeSafe has published in a week.

  15. Sep 21

    Jev Ultrafast: Browser Use Puts Jev in the Browser Agent Loop

    Browser Use, the team behind the popular browser-use agent library, shipped Jev Ultrafast — an open-source browser agent that reads a structured element table instead of screenshots and lets Jev pick an operation and a target element per step, with a small LLM only invoked to write text. The published demo completes a real Google Flights search in 7.1 seconds, with independent outcome verification.

  16. Sep 21

    Kev's Real Numbers: Inside the Open-Source Jev Clone's 0.8B/4B/9B Family

    explainx.ai previously covered Kev only through unverified digest headlines — a 0.5B MacBook model, then an "8B" follow-up with no source. The project's actual GitHub README and Hacker News launch thread are now public, and they describe something different: a documented 0.8B/4B/9B family with LoRA adapters, a pointer-head architecture, and benchmark numbers run against Jev on both trained and unseen data.

  17. Sep 21

    Laya-MLX: A Real On-Device Alternative to Jev — Is It Really 50x Faster?

    Developer mizorewww's laya-mlx is a real, open-source Apache-2.0 MLX port of Convai Innovations' Laya typed-decision model, with published benchmarks showing sub-14ms decisions and under 1GB peak memory on an M3 Max. A viral Chinese-language X post calls it "50x faster than Jev" — a claim the project's own README never makes. Here's what's actually measured, what isn't, and how it fits next to TypeSafe AI's Jev.

  18. Sep 21

    TypeSafe's Founder Published Coding-Agent Notes. The KV-Cache Math Is the Part Worth Reading.

    TypeSafe AI founder Diogo Almeida published a long, explicitly speculative notes document on what a Jev-centric coding agent could look like — and hopes the community builds it before he does. The most concrete, checkable claim inside is a worked cost comparison showing that routing a task to a cheaper model and back to a stronger one can cost more than never switching, because the stronger model has to reprocess the whole context from scratch.

  19. Sep 20

    Awesome Jev Use Cases: A 50-Demo Gallery You Can Run Yourself

    Every Jev use-case argument so far has been reasoning about the shape of the Choice, Score, and Noul primitives. The awesome-jev-use-cases repo skips the reasoning and ships 50 runnable demos instead — each one a side-by-side comparison against OpenAI's Responses API with a live 2D visualization, no API key needed until you want your own numbers.

  20. Sep 20

    Bespoke Nimble: A 9B Model Hit 90% on Jev — Built in Days, Not Months

    A small, independently-built open model reportedly closed most of the gap to TypeSafe AI's Jev on its own eval, in a fraction of the time a frontier lab spends on a launch. The real story isn't the score — it's what a 9B LoRA fine-tune beating expectations says about efficient fine-tuning versus brute-force scale, and how skeptically to read any "hits X% on eval Y" headline.

  21. Sep 20

    Jev vs LLM-as-Judge: LangChain Benchmarks Agent Evaluation

    LangChain ran the same Deep Agents weather-tool traces through four judges — TypeSafe AI's Jev, GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 — and measured accuracy against a human oracle, per-case variance, cost, and latency. Jev matched the human oracle on all 500 repeated decisions at roughly 1/80,000th the cost of Claude.

  22. Sep 20

    Top 10 Ways to Learn Jev (TypeSafe AI) in 2026: Courses, Workshops, and Resources

    Jev, TypeSafe AI's "System One Model," is five days old and search results for "Jev course" are already a mess of speculation. Here's an honest, ranked list of the real resources worth your time — starting with explainx.ai's own self-paced Udemy course and live Jev workshop.

  23. Sep 19

    How to Wire Jev Into Your Agent Pipeline for Routing Decisions

    Jev is available directly on Vercel's AI Gateway, exposed through AI SDK 7's experimental_evaluate function, and has an official LangChain integration (TypeSafeClassifier) built specifically for routing, escalation, and tool-call decisions inside an agent loop. Here's how to actually wire it in, with the concrete integration points and what each one is for.

  24. Sep 19

    Jev's Actual Security Use Case: Detecting Prompt Injection, Not Getting Hacked

    There's no published adversarial research on gaming or poisoning Jev, TypeSafe AI's non-generative "System One Model" — a search for that angle comes up thin. What does exist is the inverse: Jev being positioned as a security tool itself, with a `contains_prompt_injection` classification primitive meant to sit in front of a main LLM and flag jailbreak or injection attempts fast and cheap, before they reach the model actually generating your response.

  25. Sep 19

    Is Jev's 200x-Faster, 400x-Cheaper Claim Actually True?

    TypeSafe AI's headline numbers for Jev — 20-200x faster, 40-400x cheaper than LLMs on structured-output tasks — are TypeSafe's own benchmarks, measured against agreement with other frontier models rather than verified ground truth. An independent test from Every corroborated the general direction but called results "good but not perfect," and Jev's own dashboard shows a real accuracy gap against the best comparator model.

  26. Sep 19

    Jev vs. XGBoost and BERT: Is a System One Model Actually New?

    Before Jev, teams needing fast structured classification typically reached for XGBoost (fast, cheap, lower ceiling on accuracy) or a fine-tuned BERT model (higher accuracy, more setup, still not free-text generation). Jev sits in a genuinely different spot on that spectrum — not because typed-output classification is new, but because of how it's trained and how it reports confidence. Here's an honest comparison.

  27. Sep 19

    OpenJev: Try Jev's Trick Yourself, Free, in Your Browser

    A developer built OpenJev — a free, entirely in-browser tool that lets anyone run open models like Qwen3 and MiniCPM5 locally and directly compare TypeSafe's Jev-style "direct readout" decision method against ordinary token-by-token generation, on their own GPU. It hit 556 points on Hacker News, and the discussion is as much about a naming dispute and a vibecoded-looking UI as it is about the underlying technique.

  28. Sep 19

    Six Jev Clones Shipped in Two Days — Here's What Each One Actually Does

    TypeSafe AI's Jev launched September 15, 2026. Within two days, at least six independent open-source clones or alternatives appeared, catalogued by Latent.Space — ranging from a 421M-parameter ModernBERT-based model to a 40KB embedding-only implementation to a 0.5B model designed to run on a MacBook Pro. Here's what each one actually is, and what the speed of the response says about how replicable Jev's core idea turned out to be.

  29. Sep 19

    Where Jev Actually Fails: The Specific Complaints Behind the Hype

    Jev's Hacker News launch thread ran to 256 comments, and buried in the general skepticism are specific, concrete failure modes worth taking seriously — not "it's not an LLM" complaints, but named cases where Jev returns a type-valid, well-formed, confidently-scored answer that is simply wrong. Here's what's actually been reported, sourced directly.

  30. Sep 17

    Could Jev Run Self-Driving? The Latency Argument, and Why Engineers Pushed Back

    "The implications of Jev on self-driving could be huge" drew 209,000 views and an immediate wall of pushback from people who work on autonomy. Their objections are specific and they are right, but the underlying question of where a fast decision model belongs in a robotics stack is still a good one.

Other topics

  • Claude Code
  • OpenAI Codex
  • AI Coding Tools
  • Model Context Protocol (MCP)
  • Agent Skills
  • Anthropic and Claude
  • OpenAI and ChatGPT
  • Google Gemini and DeepMind
  • Meta AI
  • xAI and Grok
  • Microsoft, Apple and Amazon AI
  • Open-Weight Models
  • Local AI
  • AI Agents
  • AI Safety and Alignment
  • AI Policy and Regulation
  • AI Security
  • AI Chips and Infrastructure
  • Robotics and Physical AI
  • AI Benchmarks and Evals
  • AI Research
  • AI Image, Video and Voice
  • Prompt Engineering
  • Learning AI and Careers
  • AI Tools and Apps
  • AI Industry and Business