explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: What people are asking
  • What Jeff actually is (repo, not digest)
  • Is this just structured outputs / constrained decoding / logprobs?
  • vs BERT, FLAN-T5, ModernBERT, embeddings + logistic
  • Jeff vs hosted Jev vs Kev vs Laya vs OpenJev vs AutoJev
  • The benchmark table you should actually read
  • Why 0.8B beats 2B, wording, zero-shot vs fine-tune
  • Price and speed versus DeepSeek Flash (and hosted Jev)
  • Honest HN: 70% vs 94%, job ads, "they all suck"
  • How to call the Jev-compatible API locally
  • What this means for builders
  • Related reading
← Back to blog

explainx / blog

Jeff: Home-Trained Jev-Compatible 0.8B Decision Models

Jev, Jeff, Open Source AI, Qwen3.5, Decision Models, Local AI

Jeff is firelex's Jev-compatible 0.8B/2B local decision model: 28ms on M4 Max, Apache-2.0 weights. HN asked if it's just logprobs — the repo's answer.

Sep 29, 2026·12 min read·Yash Thakker
add explainx.ai
go deep
Jeff: Home-Trained Jev-Compatible 0.8B Decision Models

Jeff hit the Hacker News front page on 28 September 2026 as "Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms" (item 49883844, 289 points and 117 comments when we pulled the thread on 29 September 2026). The linked repo is firelex/jeff — 363 stars and 9 forks on that same morning, created 28 September 2026. This is not another write-up of Kev, OpenJev, or the two-day clone wave. It is the home-trained Qwen3.5 / Gemma checkpoint family that speaks TypeSafe Jev's request format without being TypeSafe.

The author's own launch comment is the right briefing: open-weight Qwen3.5 and Gemma fine-tunes for zero-shot classification; calibrated probabilities in one forward pass; no text generation; Apache 2.0 weights; Jev-compatible API; not affiliated with TypeSafe. Training ran on one RTX PRO 6000, synthetic data from Qwen3.8-Flash-Next on two DGX Sparks, tests on a MacBook. explainx.ai checked the GitHub README and Hugging Face cards against that HN paste. Where they agree, we cite the repo. Where HN rounded, we use the README.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: What people are asking

table · 2 cols
QuestionDirect answer
Is this just structured outputs / constrained decoding / logprobs?No for JSON-mode and constrained decoding. Not "just" logprobs either: trained readout + full-weight fine-tune + fitted temperature.
vs BERT / FLAN-T5 / ModernBERT / embeddings+logistic?Those are usually fixed-label classifiers (or a separate encoder+head). Jeff's options are described at request time (up to 255). Embeddings+logistic can still win on labeled, high-class-count tasks — HN said so.
vs hosted Jev vs Kev vs Laya vs OpenJev?Jev = cloud API. Jeff = local Jev-shaped HTTP. Kev = Palmer LoRA family. Laya = encoder typed-decision + MLX port. OpenJev = in-browser readout vs generation.
Why does 0.8B beat 2B?Author: 2B more risk-averse; training amplified it. Panel accuracy still favors 2B.
Zero-shot vs fine-tune?Voice-nav held-out: 31.7% → 95.8% in under ~30 minutes on one GPU (autojev-train --initial-checkpoint).
Price/speed vs DeepSeek Flash?Jeff has no token invoice once weights are local (~22 ms RTX PRO 6000, 28 ms M4 Max MLX for 0.8B). Flash is a generative API you pay per token. Do not compare them as the same product.
Honest 70% vs 94% / job ads?Real HN reports. Treat panel 83.1% as not your job-ad accuracy.

What Jeff actually is (repo, not digest)

Project jeff version 0.2.0 (pyproject). GitHub description: "Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification." Hugging Face cards (publisher mstrasser):

table · 5 cols
CheckpointBasePanel accuracy (Jeff)ECE~16-bit size
Jeff-Qwen3.5-0.8BQwen/Qwen3.5-0.8B79.1%0.0491.7 GB
Jeff-Qwen3.5-2BQwen/Qwen3.5-2B83.1%0.0284.2 GB
Jeff-Gemma4-E2Bgoogle/gemma-4-E2B-it81.6%0.0319.3 GB stored (2B effective)

Jev's published five-benchmark overall is 83.0%. The README is explicit: Jev and AutoJev figures were measured on a different sample of the same benchmarks. Do not treat 83.1 vs 83.0 as a controlled win.

Training recipe (README): full-weight supervised fine-tuning, one epoch, batch 256, cross-entropy over option letters, then one fitted temperature for calibration. Checkpoints are chosen on a development set, never on the published panel. That is a different stack from Kev's LoRA adapter and pointer head. Code is MIT (copyright Mathias Strasser / Jeff, plus Denis Yarats / AutoJev). Model weights: Apache 2.0. Training data is not released; sources and licences live in the repo's docs/data-sources.md.

A System One Model here means choice / noul / score over enumerated options, not a chat LLM.

Open-weight versus closed-model choice as a metaphor for running Jeff locally instead of a hosted Jev API

Is this just structured outputs / constrained decoding / logprobs?

Structured outputs / JSON mode: no. Jeff's README: "No generated text, no parsing." The server exposes POST /v1/systemone, not a chat completions grammar.

Constrained decoding: no. There is no token-by-token sampler constrained to a JSON schema. The model scores option codes in one forward pass.

Logprobs on a frozen chat model: closer, still incomplete. The AutoJev/Jeff design uses a trained answer readout over single-token option codes (initialized from output embeddings in the generic decoder path). You are reading a distribution over a small label set. You are not wrapping GPT-style generation and scraping logprobs. Weights are fine-tuned; a temperature is fitted for calibration. Calling that "just logprobs" is like calling a fine-tuned BERT "just softmax." Related explainx.ai frame: Jev vs XGBoost and BERT.

vs BERT, FLAN-T5, ModernBERT, embeddings + logistic

HN commenter nico (replying to the 70%-vs-94% thread) reported embeddings plus a logistic head matching or beating Jev and Laya on AG News, Emotion, MASSIVE Intent, and Banking77, especially on taxonomies with more than 50 classes, at under 1 ms, and failing on reasoning-ish sets like XNLI where Jev/Laya look stronger.

That is the honest split:

  • Labeled, stable taxonomy (job industry, Banking77): a dedicated classifier or encoder often wins on accuracy and latency. Jeff's 0.8B failing job-ad fields is consistent with that.
  • Runtime-described options (this ticket's three queues, this turn's legal Doom moves): you need something that does not freeze a softmax width at train time. That is the System One pitch Jeff, Jev, and Kev share.
  • ModernBERT / Laya: encoder, PPO-trained Laya is a different base than Qwen3.5 decoder fine-tunes. See the Laya-MLX write-up for Apple-silicon encoder numbers, not Jeff's Qwen MLX path.
  • FLAN-T5: still typically generates (even if short). Jeff is trained not to decode a continuation.

Jeff vs hosted Jev vs Kev vs Laya vs OpenJev vs AutoJev

table · 3 cols
ProjectWhat it isWhat Jeff is not
Hosted JevTypeSafe cloud System One API, undisclosed size, RLCD brandingJeff is local weights + compatible shape
JeffHome-trained 0.8B/2B/Gemma-E2B, AutoJev-derived SFTNot TypeSafe; not 27B AutoJev
AutoJevOpen recipe (Denis Yarats) fine-tuning Qwen3.8-27B; Jeff README cites published AutoJev-27B panel 84.9%Jeff kept the recipe and shrunk the student
KevPalmer 0.8B/4B/9B Qwen3.5, LoRA + pointer head, kev.train --init_fromDifferent training; different sizes
Laya / laya-mlxModernBERT-scale typed decisions; MLX port claims sub-14 ms on M3 MaxDifferent architecture; Jeff's 0.8B is 28 ms on M4 Max in Jeff's table
OpenJevBrowser WebGPU, compare readout vs generationNot a trained Jeff checkpoint

Speed table from the README (median over 200 ~200-token questions, one at a time, text to probabilities):

table · 3 cols
ModelNVIDIA RTX PRO 6000Apple M4 Max (MLX)
Jeff-Qwen3.5-0.8B22 ms28 ms
Jeff-Qwen3.5-2B24 ms60 ms
Jeff-Gemma4-E2B29 ms— (MLX is Qwen-only in this repo)
Jev (published Doom)114–212 ms per API call including network

Those Jev times are not the same hardware as the Mac numbers. Jev's 200x/400x marketing is a separate fact-check; Jeff does not inherit those multipliers.

The benchmark table you should actually read

4,599 questions across five public sets, plus JevBench hard (105 items, scored separately). Numbers are the project's.

table · 6 cols
BenchmarkJeff 0.8BJeff 2BJeff Gemma E2BJev publishedAutoJev-27B published
Overall (5)79.183.181.683.084.9
BBH64.068.066.494.382.8
Financial PhraseBank96.496.396.177.084.2
JudgeBench62.664.660.678.678.9
RAGTruth86.188.987.477.388.9
WinoGrande68.679.077.490.783.3
JevBench hard47.653.348.673.370.3

HN's "BBH 64–68% vs Jev 94%" and "JevBench hard ~50% vs 73%" match this table (0.8B hard is 47.6%, 2B 53.3%). Overall closeness is classification and grounding; reasoning is where the small students lose, which is what you should expect from 0.8B–2B versus a larger hosted model. For how to read self-tables, see how to read AI benchmarks.

Why 0.8B beats 2B, wording, zero-shot vs fine-tune

Games (20 episodes, seed 1234), README:

table · 4 cols
ModelDoom killsFrogger crossingsPac-Man pellets / 98
Hand-coded rule bot6.5510.2594.1
Jeff-Qwen3.5-0.8B6.5510.357.0
Jeff-Qwen3.5-2B−0.96.041.2
Untrained Qwen3.5-0.8B5.01.025.8
Jev published Doom6.55 with aiming rule; −0.60 without——

The 2B is worse than the 0.8B on Doom and Pac-Man after training. The author: bigger is not better for "fast option picking"; the 2B is "more cautious." Untrained Gemma 4 E2B beats untrained Qwens on the panel (62.5%) and plays games worst.

Wording: giving Frogger's goal step the same language as other forward options moved one episode from 15 to 23 crossings. Jev's Doom prompt (bearing number + aiming rule) does not work for Jeff; consequence-in-words options do.

Fine-tune: voice navigation, ~11k app-specific examples, ~half an hour, one GPU, held-out 31.7% → 95.8%, ~40 ms on M4 Max. Command shape in the README: autojev-train --initial-checkpoint pointing at a Jeff checkpoint, --epochs 1. Zero-shot is the demo; domain data is the production path — same lesson as Kev's --init_from, different trainer.

Price and speed versus DeepSeek Flash (and hosted Jev)

Jeff: electricity + disk. Optional JEFF_API_KEY on the local server. TypeSafe's list price in explainx.ai's Jev coverage is $42 per billion input tokens, output free — a usage meter. DeepSeek Flash is a full generative API (explainx.ai has covered its token-volume and coding-benchmark stories separately). You do not "replace Flash" with Jeff unless your workload is typed decisions, not code/chat reasoning.

Rough practitioner math: a 200-token Jeff decision at 28 ms locally has zero Flash/Jev invoice. A Flash call that writes a paragraph is a different SKU. If you were using Flash as a classifier, Jeff (or BERT, or embeddings+logistic) is the category to A/B, not "is Flash faster than 28 ms" without matching the task.

HN asked for a "price comparison for the masses." The accurate answer is: Jeff's marginal cost is your GPU/CPU; Jev's is TypeSafe's meter; Flash's is DeepSeek's meter and generation length.

Honest HN: 70% vs 94%, job ads, "they all suck"

Do not sand this down.

  • AgentMasterRace: compared Jeff to Jev on current use cases — "70% vs 94%. for classification, it's unacceptable."
  • Oras: job ads (industry, remote/hybrid/onsite, full-time/part-time). Side-by-side Gemini 2.5 Flash Lite, Jev, Jeff. "0.8B model, completely useless." Jeff-Qwen3.5-2B better, still missed job type.
  • zergrush: "i've tried all the open source me too Jevs they all suck."
  • tbeseda: the point is an open-weight MVP you can fine-tune, not displacing Jev on day one.
  • senko: ModernBERT ~0.4B already fine-tunes; Jev's value is zero-shot without fine-tuning.
  • clhodapp: needing fine-tuning changes the product category.
  • ijustlovemath / neuronexmachina: "isn't Jev just a less nuanced classifier?" — "zero-shot is what makes these different."

The README already said small models "won't match Jev" on reasoning and that you should fine-tune if zero-shot is not enough. The thread is that warning with production labels attached. If you wire Jev into agent routing, do not drop in Jeff-0.8B without the same eval set.

How to call the Jev-compatible API locally

From the official README (NVIDIA/CPU). Apple silicon: uv sync --extra mac and JEFF_BACKEND=mlx.

bash
uv sync
uv run hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b

JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve
bash
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "model": "jeff-latest",
  "state": "Refund request: the customer says the parcel arrived crushed and wants their money back.",
  "questions": {
    "route": {"type": "choice", "instructions": "Which team should handle this?",
              "criteria": {"1": "Refunds and payments", "2": "Damaged or lost parcels", "3": "Account and login problems"}},
    "angry": {"type": "noul", "instructions": "Is the customer angry?"}
  }
}'

Question types: choice (up to 255 options), noul (yes/no as a probability), score (a scale you describe). Several independent questions in one request. README ops advice: reason in code, decide with Jeff; short option keys; describe consequences, not forecasts ("a car arrives in 2 turns" is random-level).

What this means for builders

If you need local, open-weight, Jev-shaped HTTP and you can measure your own labels, Jeff is a real checkpoint family with an unusually honest README (games vs panel, 2B vs 0.8B, wording). If you need hosted calibration and zero-shot on hard reasoning, Jev is still the product HN defenders called "king." If your taxonomy is fixed and labeled, try the classifier you already know before a 0.8B decoder.

Jeff does not replace the six clones from launch week; it is a later, AutoJev-student, train-at-home data point in the same category.

Related reading

  • Kev: open-source Jev clone, Qwen3.5 0.8B/4B/9B
  • TypeSafe AI launches Jev
  • OpenJev: browser Jev clone
  • Laya-MLX on-device alternative
  • Six Jev clones in two days
  • What is a System One Model?
  • Jev vs XGBoost and BERT
  • Jev speed/cost claims fact-check
  • Official: firelex/jeff · Hugging Face mstrasser/Jeff-Qwen3.5-0.8B · HN item 49883844

Figures, licences, and API snippets are from firelex/jeff and the mstrasser Hugging Face cards as of 29 September 2026 (repo created 28 September 2026). Hacker News score 289 / 117 comments and GitHub 363 stars were live API reads that morning — they will move. explainx.ai did not retrain Jeff or re-run the 4,599-question panel.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 21, 2026

Kev's Real Numbers: Inside the Open-Source Jev Clone's 0.8B/4B/9B Family

explainx.ai previously covered Kev only through unverified digest headlines — a 0.5B MacBook model, then an "8B" follow-up with no source. The project's actual GitHub README and Hacker News launch thread are now public, and they describe something different: a documented 0.8B/4B/9B family with LoRA adapters, a pointer-head architecture, and benchmark numbers run against Jev on both trained and unseen data.

Sep 26, 2026

Ollaya Is "Ollama for Decision Models" — A Local Runtime, Not a New Model

Every Jev clone so far has shipped a single model. Ollaya ships none of its own — instead it's a desktop app, CLI, and Docker image that bundles seven open decision models behind a drop-in TypeSafe-compatible local endpoint, the same category move Ollama made for local LLMs.

Sep 22, 2026

SemIf on LangSmith Gateway: Free Decision AI Through Sept 28

LangChain added a Decision models category to the LangSmith LLM Gateway on September 22, 2026, with hosted SemIf (semif-qwen3.5-4b) free through September 28 on US Free, Developer, and Plus workspaces. Here's what SemIf is, how its authored144 benchmark fits the Jev ecosystem, and what you get versus bringing your own TypeSafe (Jev) key.