explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • What people are asking after the headline
  • How /v1/systemone works in practice
  • Local decision-model runners: where llama.cpp sits now
  • What this changes for agent builders
  • Honest limitations
  • Quick start checklist
  • Related on explainx.ai
← Back to blog

explainx / blog

llama.cpp Adds Decision Model Support for Five Open Models

llama.cpp, Decision Models, Local AI, Open Source AI, Jev

llama.cpp PR #29818 adds /v1/systemone for Jev-style decision models. Five open GGUFs: Julia-1, Laya, Kev-4B, lev, OpenJev.

Oct 3, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
llama.cpp Adds Decision Model Support for Five Open Models

October 3, 2026: The feed headline was right — llama.cpp can run decision models now. On October 2, 2026, Xuan-Son Nguyen and Victor Mustar published the ggml-org announcement: llama-server gained POST /v1/systemone, implemented in PR #29818. You send a state (text, JSON, or a screenshot for vision-capable weights) plus typed questions. The model returns a probability for each option in one forward pass, with output_tokens: 0 in the usage block.

The announcement names exactly five open models with GGUF builds today: Julia-1, Laya, Kev-4B, lev, and OpenJev. That is the list we verified against the primary source — not a rewritten aggregator count.

This post is the practitioner read: what shipped, which models are real, how the API works, and how this path sits next to Ollaya, GLiNER2.5-Decide, Jeff, and the earlier Julia / Laya / Kev coverage on explainx.ai.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer
What shipped?Decision-model serving in llama-server via /v1/systemone (PR #29818)
When?Announced October 2, 2026 on the ggml Hugging Face blog
How many models named?Five — Julia-1, Laya, Kev-4B, lev, OpenJev
API shape?System One: state + questions (choice, score, noul)
Images?OpenJev only in the announced table (vision projector auto-downloads)
Licenses?Julia-1, Laya, Kev-4B, lev: Apache 2.0; OpenJev: CC BY-NC 4.0
vs Ollaya?Same decision-API idea; llama.cpp is the GGUF engine you may already run
vs GLiNER / Jeff?Different stacks — Python encoder (GLiNER) and Firelex server (Jeff), not this PR
Do I need a new product?No if you already update llama.cpp — llama serve -hf ggml-org/...

What people are asking after the headline

Can llama.cpp actually run Jev-style classifiers now?

Yes — for open weights that the ggml team converted and wired into the decision path. A decision model is not a chat LLM with JSON mode. It scores the options you provide instead of generating text. The ggml post puts it plainly: a chat model spends one forward pass per output token and still needs parsing; a decision model reads the input once and returns one of your options with a probability attached.

Typical jobs they name: routing a request, moderating content, checking that an agent step worked, or choosing the next action. That matches the category explainx.ai has been covering since TypeSafe's Jev launch — see also What is a System One model?.

Which five models — and did we invent any names?

We did not. The official supported-models table lists:

table · 7 cols
ModelSizeBased onLanguagesImagesLicenseSpeed*
Julia-1144MmmBERT-small50+noApache 2.0~3 ms
Laya421MModernBERT-largeEnglishnoApache 2.0~5 ms
Kev-4B4BQwen3.5-4B-BaseEnglishnoApache 2.0~12 ms
lev4BQwen3.5-4BEnglishnoApache 2.0~36 ms
OpenJev27BQwen3.8-27Ben, de, fr, hi, zh, jayesCC BY-NC 4.0~43 ms

*Median time to answer one question on one NVIDIA RTX PRO 6000, per the announcement.

explainx.ai already covered several of these outside llama.cpp packaging: Julia 1 on CPU, Laya on MLX, Kev's Qwen3.5 family, and OpenJev as a browser / open-model demo. What is new is a first-party GGUF + llama-server path with the System One endpoint baked in — not a third-party fork claiming compatibility.

lev appears in the ggml table as a fourth English 4B-class option (alongside Kev-4B) based on Qwen3.5-4B. Treat the announcement as the naming authority; do not assume it is the same checkpoint as every community "Jev clone" with a similar name.

Is hosted Jev now "open" through llama.cpp?

No. TypeSafe's hosted Jev remains a cloud API. OpenJev is an open-weight (non-commercial license) decision model in the same API family. The llama.cpp path runs GGUF decision models that speak the System One request shape. Pointing your client at localhost:8080 does not download Jev's private weights.

How /v1/systemone works in practice

Update to a current llama.cpp build (llama update or install from the project's current binary channel), then start a model from Hugging Face:

bash
llama serve -hf ggml-org/Kev-4B-GGUF

Three question types ship in the announcement:

table · 3 cols
TypeYou sendYou get
choiceoptions, optional descriptionstop option + probability per option
score2 to 10 levels, lowest firstexpected level (can fall between two)
noula yes/no questionprobability of yes

A minimal curl shape (abbreviated from the official example):

bash
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Customer message: I was charged twice for my order last week.",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": "payments, charges, refunds, invoices",
          "shipping": "delivery, tracking, lost or late parcels",
          "technical": "bugs, errors, login problems"
        }
      },
      "angry": {
        "type": "noul",
        "instructions": "Is the customer angry?"
      },
      "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["can wait", "this week", "today", "right now"]
      }
    }
  }'

The sample response returns structured answers with probabilities, optional confidence, and usage showing input tokens only. That is the operational difference versus prompting a chat model for JSON: there is nothing to parse out of free text, and there is no token-by-token decode loop for the label.

Images and router mode

OpenJev is the vision-capable entry in the announced set. The vision projector downloads automatically when you serve ggml-org/OpenJev-GGUF. You can pass base64 data-URL images (or chat-message image_url parts) for document / screenshot classification.

Router mode still applies: run bare llama serve, then pick a model id per request (for example ggml-org/Julia-1-GGUF:Q8_0). /v1/models lists ids. With a single model loaded, the model field is ignored.

Tips that matter more than the marketing table

The announcement's practical tips are worth treating as product requirements:

  • Describe your options. Bare labels misroute; short descriptions fix Julia-1 in their example (shipping vs billing on a double-charge ticket).
  • Calibrate confidence cutoffs per model. The same vague ticket scored about 0.25 confidence on Julia-1 and about 0.80 on Kev-4B in their write-up. Thresholds do not transfer across sizes.
  • Batch questions. Kev-4B, lev, and OpenJev process the state once for multiple questions.
  • Try quants. Same GGUF story as chat models — e.g. llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0.

For the broader engine (Metal flags, -ngl, OpenAI chat on :8080), stay on the llama.cpp foundation guide. Decision serving is a new endpoint on the same binary, not a separate install.

Local decision-model runners: where llama.cpp sits now

Until this week, explainx.ai's local map looked roughly like this:

  1. Hosted System One — TypeSafe Jev (and later hosts that speak the same contract).
  2. Dedicated local runtimes — Ollaya as "Ollama for decision models," bundling several open classifiers behind /v1/systemone on its own port.
  3. Model-specific stacks — Julia 1 Python/ONNX, Laya MLX, Kev training/serve paths, Jeff with jeff-serve, GLiNER2.5-Decide via pip install gliner2.
  4. Hack paths — grammar / logprob tricks on chat models when you refuse a dedicated head.

llama.cpp is now a first-class #2.5: same System One HTTP contract, but through the GGUF engine that already powers a huge share of local chat and agent setups. If your laptop already has llama-server for OpenCode or a coding harness, adding a decision hop is an update + a -hf pull, not a new product category install.

llama.cpp vs Ollaya

table · 3 cols
Dimensionllama.cpp (this ship)Ollaya
What it isGeneral inference engine + new decision endpointDecision-model runtime / catalogue
API/v1/systemone on llama-server/v1/systemone (TypeSafe-compatible)
PackagingGGUF via ggml-org HF reposBundled model families behind one CLI/app
Best ifYou already standardize on llama.cppYou want a decision-only install with multiple families pre-wired
OverlapJulia / Laya / Kev-class models appear in both ecosystems under related namesBroader catalogue (NLI, GLiClass, guards, etc.) as of September coverage

Neither "wins" in the abstract. Ollaya optimized for the decision category. llama.cpp optimized for one binary that already does chat, embeddings, tools, and now typed decisions. Teams that hate dependency sprawl will prefer the engine they already pin. Teams that want a decision-model appliance will keep Ollaya.

llama.cpp vs GLiNER2.5-Decide and Jeff

GLiNER2.5-Decide is a 340M open-weight encoder path with its own Python API, span/relation extras, and Fastino packaging. It is not one of the five GGUFs in the ggml announcement. If your pipeline is already Python-native and you care about constraint decoding across related fields, stay on that stack until someone ships a GLiNER GGUF into llama.cpp (the announcement says more models are coming; Cloudflare's Clef is named as next).

Jeff is Firelex's home-trained 0.8B/2B line with its own local server. Also not in the five-name ggml table. Jeff remains a useful small-model experiment; llama.cpp's Julia-1 / Laya entries cover a similar "tiny local classifier" niche with official GGUF conversion.

vs Julia / Laya / Kev posts you already read

Those posts still matter for training origin, accuracy caveats, and non-GGUF runtimes. This news post answers a different question: can the default local LLM runner serve them without a bespoke server? For the five named GGUFs, the answer is yes on current llama.cpp builds.

What this changes for agent builders

Decision models earn their keep when an agent makes the same closed-set call thousands of times — route, gate, score, moderate — and you want a calibrated probability you can threshold. Wiring that hop through llama.cpp means:

  • One process family for chat completions and System One decisions (different endpoints, same server lifecycle).
  • Quantization and backend flags you already know — Metal layers, CUDA offload, Q4/Q8 tradeoffs — apply to the decision GGUFs.
  • Router mode can park Julia-1 for cheap gates and Kev-4B / OpenJev for harder English or multilingual / vision hops without standing up three products.

It does not mean accuracy equals hosted Jev. explainx.ai has repeatedly documented gaps on hard classification and reasoning-shaped tasks for small open clones. Run your own labeled set. Pick cutoffs per model. Escalate low-confidence rows to a larger decision model or a human — the announcement itself recommends that pattern.

Also respect licenses: OpenJev's CC BY-NC 4.0 blocks commercial use even if the weights download fine. The Apache-2.0 quartet (Julia-1, Laya, Kev-4B, lev) is the safer commercial shortlist until counsel reviews OpenJev.

Honest limitations

  • Not every decision model on the internet is supported. Adding a new architecture needs a DecisionType in the conversion path and wiring in server-decision.cpp (documented in the project's HOWTO-add-model notes pointing at PR #29818). Community clones without a GGUF conversion stay on their own servers.
  • GLiNER2.5-Decide, Jeff, and some Ollaya catalogue entries are not in the five-name table. Do not assume the headline means "every Jev clone now runs in llama.cpp."
  • Speed numbers are vendor-hardware. The ~3–43 ms table is RTX PRO 6000 medians. Your M-series laptop or CPU box will differ.
  • Confidence is not a unit test. Same warning as the category guide: repeated calls can differ; schema validity is not determinism.
  • Clef is "next," not here. The announcement teases Cloudflare's Clef as a forthcoming addition. Until it lands in a build you can pin, do not plan production on it inside llama.cpp.

Quick start checklist

  1. Update llama.cpp to a build that includes PR #29818's decision path.
  2. Pick a license-appropriate GGUF — start with ggml-org/Julia-1-GGUF or ggml-org/Kev-4B-GGUF for English text.
  3. llama serve -hf ggml-org/Kev-4B-GGUF (add a quant tag if you want).
  4. POST to http://localhost:8080/v1/systemone with state + questions.
  5. Log confidence and option probabilities; set escalate thresholds from your data, not from the blog table.
  6. If you need vision documents, switch to OpenJev and confirm CC BY-NC is acceptable.
  7. If you need Fastino-style constraint decoding or Jeff's specific checkpoints, keep those stacks — this PR does not subsume them.

Related on explainx.ai

  • What is llama.cpp? — install, GGUF, llama-server basics
  • What are decision models? — category architecture and when to use them
  • Ollaya local decision runtime — dedicated "Ollama for decision models"
  • GLiNER2.5-Decide — 340M open-weight Fastino path
  • Jeff 0.8B decision models — Firelex local Jev-compatible fine-tunes
  • Julia 1 CPU decision model · Laya on MLX · Kev family

Primary source: ggml-org Hugging Face blog post "New in llama.cpp: Decision Models" (October 2, 2026) and llama.cpp PR #29818.

Specs, model lists, licenses, and latency figures are accurate as of October 3, 2026 against that announcement. Builds move fast — pin a bNNNN release and re-check /v1/models before production.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 29, 2026

Jeff: Home-Trained Jev-Compatible 0.8B Decision Models

A Hacker News front-page project from firelex ships Qwen3.5 and Gemma fine-tunes that speak Jev's request format locally. The 2B's five-benchmark panel score is 83.1% against Jev's published 83.0% on a different sample; the 0.8B is ~28 ms on an M4 Max. This post is about Jeff specifically — not Kev, OpenJev, or Laya — and what the thread actually argued.

Sep 26, 2026

Ollaya Is "Ollama for Decision Models" — A Local Runtime, Not a New Model

Every Jev clone so far has shipped a single model. Ollaya ships none of its own — instead it's a desktop app, CLI, and Docker image that bundles seven open decision models behind a drop-in TypeSafe-compatible local endpoint, the same category move Ollama made for local LLMs.

Sep 21, 2026

Kev's Real Numbers: Inside the Open-Source Jev Clone's 0.8B/4B/9B Family

explainx.ai previously covered Kev only through unverified digest headlines — a 0.5B MacBook model, then an "8B" follow-up with no source. The project's actual GitHub README and Hacker News launch thread are now public, and they describe something different: a documented 0.8B/4B/9B family with LoRA adapters, a pointer-head architecture, and benchmark numbers run against Jev on both trained and unseen data.