October 3, 2026: The feed headline was right — llama.cpp can run decision models now. On October 2, 2026, Xuan-Son Nguyen and Victor Mustar published the ggml-org announcement: llama-server gained POST /v1/systemone, implemented in PR #29818. You send a state (text, JSON, or a screenshot for vision-capable weights) plus typed questions. The model returns a probability for each option in one forward pass, with output_tokens: 0 in the usage block.
The announcement names exactly five open models with GGUF builds today: Julia-1, Laya, Kev-4B, lev, and OpenJev. That is the list we verified against the primary source — not a rewritten aggregator count.
This post is the practitioner read: what shipped, which models are real, how the API works, and how this path sits next to Ollaya, GLiNER2.5-Decide, Jeff, and the earlier Julia / Laya / Kev coverage on explainx.ai.
TL;DR
| Question | Answer |
|---|---|
| What shipped? | Decision-model serving in llama-server via /v1/systemone (PR #29818) |
| When? | Announced October 2, 2026 on the ggml Hugging Face blog |
| How many models named? | Five — Julia-1, Laya, Kev-4B, lev, OpenJev |
| API shape? | System One: state + questions (choice, score, noul) |
| Images? | OpenJev only in the announced table (vision projector auto-downloads) |
| Licenses? | Julia-1, Laya, Kev-4B, lev: Apache 2.0; OpenJev: CC BY-NC 4.0 |
| vs Ollaya? | Same decision-API idea; llama.cpp is the GGUF engine you may already run |
| vs GLiNER / Jeff? | Different stacks — Python encoder (GLiNER) and Firelex server (Jeff), not this PR |
| Do I need a new product? | No if you already update llama.cpp — llama serve -hf ggml-org/... |
What people are asking after the headline
Can llama.cpp actually run Jev-style classifiers now?
Yes — for open weights that the ggml team converted and wired into the decision path. A decision model is not a chat LLM with JSON mode. It scores the options you provide instead of generating text. The ggml post puts it plainly: a chat model spends one forward pass per output token and still needs parsing; a decision model reads the input once and returns one of your options with a probability attached.
Typical jobs they name: routing a request, moderating content, checking that an agent step worked, or choosing the next action. That matches the category explainx.ai has been covering since TypeSafe's Jev launch — see also What is a System One model?.
Which five models — and did we invent any names?
We did not. The official supported-models table lists:
| Model | Size | Based on | Languages | Images | License | Speed* |
|---|---|---|---|---|---|---|
| Julia-1 | 144M | mmBERT-small | 50+ | no | Apache 2.0 | ~3 ms |
| Laya | 421M | ModernBERT-large | English | no | Apache 2.0 | ~5 ms |
| Kev-4B | 4B | Qwen3.5-4B-Base | English | no | Apache 2.0 | ~12 ms |
| lev | 4B | Qwen3.5-4B | English | no | Apache 2.0 | ~36 ms |
| OpenJev | 27B | Qwen3.8-27B | en, de, fr, hi, zh, ja | yes | CC BY-NC 4.0 | ~43 ms |
*Median time to answer one question on one NVIDIA RTX PRO 6000, per the announcement.
explainx.ai already covered several of these outside llama.cpp packaging: Julia 1 on CPU, Laya on MLX, Kev's Qwen3.5 family, and OpenJev as a browser / open-model demo. What is new is a first-party GGUF + llama-server path with the System One endpoint baked in — not a third-party fork claiming compatibility.
lev appears in the ggml table as a fourth English 4B-class option (alongside Kev-4B) based on Qwen3.5-4B. Treat the announcement as the naming authority; do not assume it is the same checkpoint as every community "Jev clone" with a similar name.
Is hosted Jev now "open" through llama.cpp?
No. TypeSafe's hosted Jev remains a cloud API. OpenJev is an open-weight (non-commercial license) decision model in the same API family. The llama.cpp path runs GGUF decision models that speak the System One request shape. Pointing your client at localhost:8080 does not download Jev's private weights.
How /v1/systemone works in practice
Update to a current llama.cpp build (llama update or install from the project's current binary channel), then start a model from Hugging Face:
llama serve -hf ggml-org/Kev-4B-GGUF
Three question types ship in the announcement:
| Type | You send | You get |
|---|---|---|
choice | options, optional descriptions | top option + probability per option |
score | 2 to 10 levels, lowest first | expected level (can fall between two) |
noul | a yes/no question | probability of yes |
A minimal curl shape (abbreviated from the official example):
curl http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"state": "Customer message: I was charged twice for my order last week.",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "payments, charges, refunds, invoices",
"shipping": "delivery, tracking, lost or late parcels",
"technical": "bugs, errors, login problems"
}
},
"angry": {
"type": "noul",
"instructions": "Is the customer angry?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["can wait", "this week", "today", "right now"]
}
}
}'
The sample response returns structured answers with probabilities, optional confidence, and usage showing input tokens only. That is the operational difference versus prompting a chat model for JSON: there is nothing to parse out of free text, and there is no token-by-token decode loop for the label.
Images and router mode
OpenJev is the vision-capable entry in the announced set. The vision projector downloads automatically when you serve ggml-org/OpenJev-GGUF. You can pass base64 data-URL images (or chat-message image_url parts) for document / screenshot classification.
Router mode still applies: run bare llama serve, then pick a model id per request (for example ggml-org/Julia-1-GGUF:Q8_0). /v1/models lists ids. With a single model loaded, the model field is ignored.
Tips that matter more than the marketing table
The announcement's practical tips are worth treating as product requirements:
- Describe your options. Bare labels misroute; short descriptions fix Julia-1 in their example (shipping vs billing on a double-charge ticket).
- Calibrate confidence cutoffs per model. The same vague ticket scored about 0.25 confidence on Julia-1 and about 0.80 on Kev-4B in their write-up. Thresholds do not transfer across sizes.
- Batch questions. Kev-4B, lev, and OpenJev process the state once for multiple questions.
- Try quants. Same GGUF story as chat models — e.g.
llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0.
For the broader engine (Metal flags, -ngl, OpenAI chat on :8080), stay on the llama.cpp foundation guide. Decision serving is a new endpoint on the same binary, not a separate install.
Local decision-model runners: where llama.cpp sits now
Until this week, explainx.ai's local map looked roughly like this:
- Hosted System One — TypeSafe Jev (and later hosts that speak the same contract).
- Dedicated local runtimes — Ollaya as "Ollama for decision models," bundling several open classifiers behind
/v1/systemoneon its own port. - Model-specific stacks — Julia 1 Python/ONNX, Laya MLX, Kev training/serve paths, Jeff with
jeff-serve, GLiNER2.5-Decide viapip install gliner2. - Hack paths — grammar / logprob tricks on chat models when you refuse a dedicated head.
llama.cpp is now a first-class #2.5: same System One HTTP contract, but through the GGUF engine that already powers a huge share of local chat and agent setups. If your laptop already has llama-server for OpenCode or a coding harness, adding a decision hop is an update + a -hf pull, not a new product category install.
llama.cpp vs Ollaya
| Dimension | llama.cpp (this ship) | Ollaya |
|---|---|---|
| What it is | General inference engine + new decision endpoint | Decision-model runtime / catalogue |
| API | /v1/systemone on llama-server | /v1/systemone (TypeSafe-compatible) |
| Packaging | GGUF via ggml-org HF repos | Bundled model families behind one CLI/app |
| Best if | You already standardize on llama.cpp | You want a decision-only install with multiple families pre-wired |
| Overlap | Julia / Laya / Kev-class models appear in both ecosystems under related names | Broader catalogue (NLI, GLiClass, guards, etc.) as of September coverage |
Neither "wins" in the abstract. Ollaya optimized for the decision category. llama.cpp optimized for one binary that already does chat, embeddings, tools, and now typed decisions. Teams that hate dependency sprawl will prefer the engine they already pin. Teams that want a decision-model appliance will keep Ollaya.
llama.cpp vs GLiNER2.5-Decide and Jeff
GLiNER2.5-Decide is a 340M open-weight encoder path with its own Python API, span/relation extras, and Fastino packaging. It is not one of the five GGUFs in the ggml announcement. If your pipeline is already Python-native and you care about constraint decoding across related fields, stay on that stack until someone ships a GLiNER GGUF into llama.cpp (the announcement says more models are coming; Cloudflare's Clef is named as next).
Jeff is Firelex's home-trained 0.8B/2B line with its own local server. Also not in the five-name ggml table. Jeff remains a useful small-model experiment; llama.cpp's Julia-1 / Laya entries cover a similar "tiny local classifier" niche with official GGUF conversion.
vs Julia / Laya / Kev posts you already read
Those posts still matter for training origin, accuracy caveats, and non-GGUF runtimes. This news post answers a different question: can the default local LLM runner serve them without a bespoke server? For the five named GGUFs, the answer is yes on current llama.cpp builds.
What this changes for agent builders
Decision models earn their keep when an agent makes the same closed-set call thousands of times — route, gate, score, moderate — and you want a calibrated probability you can threshold. Wiring that hop through llama.cpp means:
- One process family for chat completions and System One decisions (different endpoints, same server lifecycle).
- Quantization and backend flags you already know — Metal layers, CUDA offload, Q4/Q8 tradeoffs — apply to the decision GGUFs.
- Router mode can park Julia-1 for cheap gates and Kev-4B / OpenJev for harder English or multilingual / vision hops without standing up three products.
It does not mean accuracy equals hosted Jev. explainx.ai has repeatedly documented gaps on hard classification and reasoning-shaped tasks for small open clones. Run your own labeled set. Pick cutoffs per model. Escalate low-confidence rows to a larger decision model or a human — the announcement itself recommends that pattern.
Also respect licenses: OpenJev's CC BY-NC 4.0 blocks commercial use even if the weights download fine. The Apache-2.0 quartet (Julia-1, Laya, Kev-4B, lev) is the safer commercial shortlist until counsel reviews OpenJev.
Honest limitations
- Not every decision model on the internet is supported. Adding a new architecture needs a
DecisionTypein the conversion path and wiring inserver-decision.cpp(documented in the project's HOWTO-add-model notes pointing at PR #29818). Community clones without a GGUF conversion stay on their own servers. - GLiNER2.5-Decide, Jeff, and some Ollaya catalogue entries are not in the five-name table. Do not assume the headline means "every Jev clone now runs in llama.cpp."
- Speed numbers are vendor-hardware. The ~3–43 ms table is RTX PRO 6000 medians. Your M-series laptop or CPU box will differ.
- Confidence is not a unit test. Same warning as the category guide: repeated calls can differ; schema validity is not determinism.
- Clef is "next," not here. The announcement teases Cloudflare's Clef as a forthcoming addition. Until it lands in a build you can pin, do not plan production on it inside llama.cpp.
Quick start checklist
- Update llama.cpp to a build that includes PR #29818's decision path.
- Pick a license-appropriate GGUF — start with
ggml-org/Julia-1-GGUForggml-org/Kev-4B-GGUFfor English text. llama serve -hf ggml-org/Kev-4B-GGUF(add a quant tag if you want).- POST to
http://localhost:8080/v1/systemonewithstate+questions. - Log
confidenceand option probabilities; set escalate thresholds from your data, not from the blog table. - If you need vision documents, switch to OpenJev and confirm CC BY-NC is acceptable.
- If you need Fastino-style constraint decoding or Jeff's specific checkpoints, keep those stacks — this PR does not subsume them.
Related on explainx.ai
- What is llama.cpp? — install, GGUF, llama-server basics
- What are decision models? — category architecture and when to use them
- Ollaya local decision runtime — dedicated "Ollama for decision models"
- GLiNER2.5-Decide — 340M open-weight Fastino path
- Jeff 0.8B decision models — Firelex local Jev-compatible fine-tunes
- Julia 1 CPU decision model · Laya on MLX · Kev family
Primary source: ggml-org Hugging Face blog post "New in llama.cpp: Decision Models" (October 2, 2026) and llama.cpp PR #29818.
Specs, model lists, licenses, and latency figures are accurate as of October 3, 2026 against that announcement. Builds move fast — pin a bNNNN release and re-check /v1/models before production.
