explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the day-one verdict
  • What developers on Hacker News are saying
  • What people are asking
  • How to call it (from the official docs)
  • When to choose Decisions API, Jev, or open weights
  • Run your own eval before switching
  • Honest limits of this verdict
  • What this means for what you build or pay
  • Related on explainx.ai
← Back to blog

explainx / blog

OpenAI Decisions API Developer Verdict: Day-One Tests vs Jev

OpenAI, Decisions API, Jev, Decision Models, Developer Feedback

Hacker News testers ran OpenAI Decisions API against Jev on day one. Price, latency, calibration, images, compliance, and when Luna is still worth it.

Oct 7, 2026·19 min read·Yash Thakker
add explainx.ai
go deep
OpenAI Decisions API Developer Verdict: Day-One Tests vs Jev

OpenAI's Decisions API went into public beta on October 6, 2026, and within about two hours developers had pointed it at their own workloads. The Hacker News thread (roughly 190 points and 80+ comments) is the first real field report. One tester ran around 600 calls against Jev and Mercury Decide, another timed 160 to 175 ms end to end, and Simon Willison shipped a plugin before the evening was over.

The short verdict from those reports: POST /v1/decisions with gpt-6-luna is cheap and quick in absolute terms, but on day one it did not beat Jev on list price, on one tester's measured cost, or on confidence for ambiguous inputs. Its strongest case is not performance. It is procurement, compliance, and not adding a vendor. Every number below is a commenter's self-report from a small, launch-day sample, not a verified benchmark.

This post does not re-explain the announcement or the endpoint table. Those are in the DevDay write-up and its public beta update. Here the question is what happened when people actually used it.

A green bead resting at the midpoint of a cream gauge arc, standing in for the confidence reading a decision model returns

TL;DR: the day-one verdict

table · 2 cols
QuestionDay-one answer
Is it cheaper than Jev?No on list price: $0.10 vs $0.042 per million input tokens, about 2.4x. One tester measured about 3.1x on a small eval. Output is free on both.
Is it faster?Disputed. One tester saw p50 346 ms and p95 860 ms via OpenRouter and called it slower than Jev. Another saw 160 to 175 ms end to end and p95 285 ms.
Is it better calibrated?Nobody knows. OpenAI makes no calibration claim in the docs we read. One tester saw lower probabilities than Jev on ambiguous tags.
Does it take images?Yes, as inline base64 only. Jev was text and JSON only at launch, which commenters called a gap.
Why pick it anyway?One vendor contract, ZDR and HIPAA for eligible customers, US and EU residency, and an SDK you already have.
Will /v1/decisions become a standard?Simon Willison thinks it might. Other vendors' schemas still differ today.
Should I switch production traffic?Not on day one. Add it as a second provider behind an adapter and run your own eval.
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What developers on Hacker News are saying

Simon Willison: a working call, a plugin, and a standards bet

simonw posted a curl call with two independent predicate questions about the sentence "I am angry about the new product feature." The API returned a probability of 0.91 for "is this a complaint" and 0.06 for "is this a compliment." His pasted usage block showed 310 input tokens and zero output tokens, which matches the docs' no-output-charge pricing.

His bigger point was about the endpoint itself. When OpenAI defines an endpoint shape like /v1/decisions, he wrote, it often becomes a de facto standard that other providers copy. That is a prediction, not a fact yet. Perplexity's decider uses choice and noul, Jev uses state plus typed questions, and llama.cpp serves /v1/systemone, as covered in our llama.cpp post. Whether those converge on OpenAI's shape is open.

He also released a plugin the same evening. We checked it on GitHub: simonw/llm-openai-decisions was created on October 6, is Apache-2.0, and has a 0.1 release on PyPI. The README documents predicates, choices, scores, multiple named questions per request, and refusals preserved as refusal answers. It accepts PNG, JPEG, WebP, and GIF attachments from a URL or local path. The source builds data: base64 URLs from those attachments, so the plugin hides the inline-base64 requirement. Install is llm install llm-openai-decisions, then llm keys set openai.

In a second comment, simonw made the point the rest of this post leans on: quality will decide the market, and anyone using a decision model has to build their own evals, because these outputs are much harder to vibe-check than text.

Topfi's roughly 600-call eval against Jev and Mercury Decide

The most detailed report came from Topfi, who ran fewer than 600 calls across UI component selection, chat charting, tag selection, and personal-knowledge-management tasks, through OpenRouter, against Jev and Mercury Decide. His own caveats are worth repeating: the evals are rudimentary, the use cases are unusual (a PKM-focused Firefox fork with infinite canvases), and he had not yet tested image input.

What he reported for Luna:

  • Latency: p50 346 ms and p95 860 ms, slower than Jev and roughly on par with Mercury Decide, though Luna's latency did not grow linearly with input size.
  • Confidence: lower confidence on his ambiguous UI-component and response-shape tasks, often at 0.6 and below.
  • Failures: four failed calls, where neither competitor had any.
  • Cost: about 3.1x Jev on average.

His summary was blunt: slower, pricier, and less capable than Jev for his workload, roughly on par with Mercury Decide, and "a bit undercooked." He added that he would rather frontier labs not jump on bandwagons without a price or performance edge. A single tester on one workload is a data point, not a ruling. With a sample this small and OpenRouter in the path, four failures cannot be separated into preview-day errors, routing-layer errors, or a systematic issue.

Latency: two testers, two pictures

agentdev001 reported a very different picture: 160 to 175 ms end to end, the worst 5 percent at 285 ms, and the worst single call at 743 ms. They also noted that because this is gpt-6-luna underneath, it should offer a 1M-token input window and image input. That window is the commenter's inference. The docs pages we read do not state a context limit.

table · 3 cols
SourceReported latencySetup disclosed
Topfip50 346 ms, p95 860 msVia OpenRouter, mixed tasks, fewer than 600 calls across three models
agentdev001160 to 175 ms typical, p95 285 ms, worst 743 msEnd to end, route and sample size not stated
OpenAI docs"10x faster than the Responses API"Relative claim, no absolute figure
OpenAI forum post"up to 10x faster than GPT-6 Luna through the Responses API"Relative claim with an "up to"

These are not necessarily contradictory. Our hypotheses, none verified: an OpenRouter hop adds time that a direct call does not; Topfi's task-specific inputs may be longer than a one-line complaint; launch-day preview load varies (Topfi himself blamed preview conditions for Mercury Decide's spikes); and neither tester published a methodology we can replicate. The Jev side is also fuzzy. TypeSafe's own materials describe a roughly 100 ms class, while the Cloudflare table cited in our decision-model category guide listed Jev near 524 ms. "Slower than Jev" depends on which Jev number you hold in your head. Another commenter, lelandfe, noted that self-hosting removes network round trips entirely, which no hosted API can match.

Price: list rate versus measured cost

On list price the answer is simple. jerrygenser put it at $0.10 per million input tokens versus $0.042 for Jev, with free output on both. Topfi followed up that in the same bench a full Jev run cost $0.0192 against $0.06 for Luna, both via OpenRouter, about 3x in Jev's favor.

The ratio deserves a second look. The list-price gap is 2.4x, but Topfi measured about 3.1x. If OpenRouter charged list rates, that would imply Luna billed roughly 30 percent more input tokens for the same eval. That is our arithmetic, not Topfi's, and it may reflect OpenAI's per-request scaffolding. Simon's 310-token usage for a roughly ten-word input hints at a fixed overhead of several hundred tokens per request, again our inference from one data point.

Scale decides whether any of this matters. At 300 tokens per call, one million decisions cost about this much:

table · 3 cols
Hosted optionInput price per 1M tokensCost of 1M decisions at 300 tokens
OpenAI Decisions API (gpt-6-luna)$0.10$30.00
Jev$0.042$12.60
Perplexity Decisions API$0.04 (a cut to $0.02 was reported by aggregators, unconfirmed)$12.00
Liquid d1$0.04$12.00
Clef-flashabout $0.09about $27
Clefabout $0.24about $72

Rates for the non-OpenAI rows come from our posts on Jev's launch, Perplexity's decider, Liquid d1, and the category guide. At 10 million decisions a month the Luna-versus-Jev gap is about $174. At 100 million it is about $1,740. For one commenter, tmhall, the whole workload comes to roughly $11 a month, and they already have OpenAI billing in place, so the cheaper vendor is not worth a new contract.

Calibration: the question nobody could answer yet

AM1010101 asked the question that matters most for anyone thresholding on these numbers: is a 90 percent output right nine times in ten? They noted that Jev had put visible effort into calibration and that it is unclear whether Luna is calibrated the same way. TypeSafe does make an explicit calibration claim for Jev's training, but our fact-check of Jev's speed and cost claims found independent verification still thin. OpenAI makes no calibration claim at all in the pages we read, only that choices and scores return a confidence and per-option probabilities.

Topfi's tag tests are the only calibration-flavored evidence so far. Each task asked whether to apply a tag to a browser tile title, with a 60 percent threshold, and every task was run twice per model:

table · 4 cols
Tile titleTagJev (two runs)Luna (two runs)
Mortgage calculator, Bankratehouse hunting0.69, 0.72 (yes)0.21, 0.21 (no)
S&P 500 index live chart, Bloomberginvesting0.91, 0.90 (yes)0.56, 0.56 (no)

Topfi's read is that Jev's values matched his own judgment better, and that Luna's output was not useful for his goal of an instant, overridable default tag. He also granted that tags are subjective and that results vary by task. Two observations from the table are ours: Luna's repeated runs matched to two decimals while Jev's drifted by a point or three, and the S&P 0.56 sits just under the 60 percent cutoff. The practical lesson is that a threshold tuned on one model does not transfer to another. A decision model's 0.56 is only meaningful against labeled examples for that model. For the broader method, see our guide to calibrating LLM classifiers.

Images: Jev's gap, not Luna's edge

mritchie712 said image input was the first big gap they found in Jev, and stillatit agreed that Decisions can take images where Jev, for now, cannot. Our own Jev launch coverage confirms text and JSON only at launch, with a 32K window.

The gap is Jev-specific rather than a Luna exclusive. Liquid d1, pplx-decider, and Clef all take images, as our posts on d1, pplx-decider, and the category guide describe. What is distinctive about OpenAI's version is the constraint: images must be inline base64 data URLs. Hosted URLs and file_id inputs are unsupported. You download, encode, and ship the bytes with every call. The docs pages we read do not say how images are tokenized, so measure usage.input_tokens on your real images before estimating cost. Liquid, by comparison, publishes a patch-based rate.

Compliance and the existing-vendor argument

The most consistent defense of Luna had nothing to do with benchmarks. Shank argued that customers with compliance needs will pick OpenAI because they cannot pick Jev. csharpminor said an enterprise with an existing OpenAI procurement agreement avoids onboarding another vendor. jcims described three- to six-month vendor onboarding cycles plus ongoing governance burden. oh_no agreed it makes sense even if underbaked, and added that it is a stopgap: stand it up, see the value, and switch in a few months if something better appears.

The official docs back the factual core. The Decisions API supports Zero Data Retention and HIPAA use "for eligible customers," with data residency and regional processing in the United States and Europe (EEA plus Switzerland). Two cautions: "eligible" depends on your agreement with OpenAI, and regional processing premiums apply on top of the $0.10 rate, so the EU price is not the headline price. We did not verify TypeSafe's compliance posture, so Shank's claim about Jev stays an assertion.

The skeptics: coherence, drift, caching, and novel domains

Coherence across questions. OutOfHere argued that decisions only suit very simple problems, and that making twenty independent calls loses coherence compared with one structured output covering many attributes. The docs half-agree: multiple independent questions can share one request, but dependent decisions need separate requests. Simon's own example shows the issue in miniature, since two independent predicates could both score high. For mutually exclusive categories, use one choice question, whose probabilities form a single distribution (the docs' sample sums to 1.0). Enforce cross-question rules in your own code. devin countered that people want more determinism, and this is a step toward it.

Silent model changes. drdexebtjl did not want a decision model to change underneath them when the lab decides to improve or "make it safer." The docs we read show one model alias, gpt-6-luna, and Simon's response echoed "model": "gpt-6-luna" with no dated snapshot visible. We could not confirm whether version pinning exists. Treat it as a question for OpenAI and defend yourself with a canary set (see below).

Caching. AM1010101 wondered why caching was skipped, since a long shared system prompt would save money on bulk work. BoorishBears guessed it chases the lowest possible latency. That guess is speculation. The docs say only that there are no cache charges.

Novel domains. nostrebored said Jev had failed to understand any novel domain they tried, and would be shocked if Luna were less generally capable. Topfi replied that results are task-dependent and his measurements had not shown Luna ahead yet. That is a genuine split, and we found no published head-to-head of Luna Decisions against Jev on a shared panel. Perplexity's 11-benchmark table shows Jev winning hard reasoning rows like BBH and WinoGrande, but it did not include Luna.

The moat debate

TSiege called OpenAI's response the nail in the coffin for the idea that AI is not a commodity market, arguing cheap System One models start a price war and the big labs now race to make products sticky. dvt saw "absolutely zero moat" and found it hard to see why anyone would pay a premium over Jev or open models. mediaman answered with the rent-versus-buy analogy: at $0.10 per million, convenience beats self-hosting even when Jev is cheaper. TSiege added that the real moat is payments, platform, and branding. scosman noted that planting a flag now and shipping a better model in a few weeks is a rational move even if today's evals are worse.

Our read: both camps are right. The vendors are separated by single-digit multiples on tiny per-call prices. The differentiators are the contract, the SDK, and the compliance paperwork.

What people are asking

Can I use my ChatGPT subscription? No. OutOfHere pointed out that ChatGPT subscriptions have never covered API calls; Decisions bills as API usage.

Is this just Luna in a wrapper? ashu1461 reasoned that the input price matches Luna's, so the difference is speed. rockinghigh added that output is free here. Per our Luna pricing post, Luna on the standard API is $0.10 per million input and $0.50 per million output, so the savings are the output tokens and the latency. Since classification outputs are short, ashu1461 countered that the output saving is small.

Will this fold into the main models? swader999 asked why routing is not simply built into every model, and jasonjmcghee wondered whether it will fold into post-training and improve calibrated outputs. Nobody has an answer.

Do I need new HTTP code? No. OutOfHere noted that the Python SDK covers it. The docs require Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0, or Java 4.78.0 and later.

How to call it (from the official docs)

This is OpenAI's own routing example, lightly trimmed. You supply model, input, and a list of uniquely named questions. A choice question takes choices, each with a value and a description.

python
from openai import OpenAI

client = OpenAI()
decision = client.decisions.create(
    model="gpt-6-luna",
    input="I was charged twice for my order.",
    questions=[
        {
            "type": "choice",
            "name": "department",
            "instructions": "Which department should handle this complaint?",
            "choices": [
                {"value": "billing", "description": "Payments, invoices, and refunds."},
                {"value": "technical", "description": "Problems using the product."},
                {"value": "shipping", "description": "Delivery and tracking."},
                {"value": "other", "description": "Requests outside these categories."},
            ],
        }
    ],
)

answer = decision.answers[0]
if answer.type == "refusal":
    print(f"Refused: {answer.name}")
elif answer.type == "choice":
    print(f"Department: {answer.choice} (confidence: {answer.confidence})")

The docs' sample response gives billing at 0.95 with an overall confidence of 0.93, plus a probability for every option. The other two types follow the same pattern. A predicate returns a probability from 0 to 1. A score question takes ordered levels and returns a probability-weighted average of the level indices, so three levels give a range of 0 to 2.

A few rules from the docs that matter more than they look:

  • Use distinct choice values with descriptions that say when each applies, and include a fallback such as other.
  • Write questions around observable criteria and separate different concerns into different questions.
  • Handle refusal answers explicitly. The docs show them as a valid answer type.
  • Send dependent decisions as separate requests.
  • For images, use input_image content with a data:image/png;base64,... URL. Nothing else is accepted.

For quick experiments from a shell, the llm plugin from the HN thread is the fastest path:

bash
llm install llm-openai-decisions
llm -m openai-decisions/gpt-6-luna 'I was charged twice.' \
  -s 'Which department should handle this message?' \
  -o answer_type choice \
  -o choices '{"billing": "Charges, invoices, and refunds", "technical": "Problems using the product", "other": null}'

For the same hop wired into an agent loop, see how to integrate Jev for agent routing and the browser version in Jev Ultrafast. The adapter pattern is the same.

When to choose Decisions API, Jev, or open weights

table · 3 cols
SituationLean towardWhy
Already on an OpenAI contract, with ZDR, HIPAA, or EU residency needsDecisions APIThe paperwork already exists, and docs list these controls for eligible customers
Image or screenshot inputsDecisions API, or another vision-capable deciderJev was text and JSON only at launch
Low volume (under about 10M calls a month)Whichever you can integrate fastestThe list-price gap is a rounding error
High volume where cost per decision dominatesJev or a cheaper hosted deciderRoughly 2.4x gap on list price, 3.1x in one measured run
You must pin a model version or keep data in your own networkOpen weights such as pplx-decider or the llama.cpp modelsYou control the checkpoint, though large ones need real GPUs
Ambiguous labels where confidence decides routingWhichever wins on your reliability plotDay-one evidence is one tester's two tag examples
UnsureBoth, behind an adapterFail over on confidence, not on brand

Open weights are not free. Perplexity's 27B decider needs roughly 49 GiB of GPU memory per our pplx-decider coverage, while the smallest llama.cpp models are in the hundreds of millions of parameters.

Run your own eval before switching

Simon Willison said anyone using a decision model has to build their own evals, and Topfi's thread is that advice in practice. Here is a minimal shape for it.

  1. Label real inputs. Collect 200 or more from your traffic, deliberately including ambiguous ones. Topfi's most informative cases were the arguable ones.
  2. Run each input at least twice per model. Topfi did, and it exposed that one model repeated to two decimals while another drifted.
  3. Record failures and refusals separately from wrong answers. A failed call is a different problem.
  4. Measure latency from your own region with enough calls for a stable p95, on more than one day. Launch-day preview load is not steady state.
  5. Log billed input tokens, especially for images, and compute dollars per thousand decisions.
  6. Fit the threshold per model on a validation split, then test on held-out data. Never reuse Jev's cutoff for Luna or the reverse.
  7. Keep a canary set of 50 to 100 frozen examples. Re-run it daily and alert when probabilities shift, which is your defense against the silent-update worry.
  8. Compare reliability, not just accuracy. Bucket predictions by probability and check whether 0.9 means about 90 percent. Our guide to reading AI benchmarks and post on evaluating prompts cover the habits.

Here is a harness that does the measuring for the predicate case. We exercised its reporting code against a stub client, but we have not pointed it at the live endpoint, so treat it as a starting point and expect to adapt field names to the SDK you install.

python
import json, statistics, sys, time
from openai import OpenAI

client = OpenAI()
QUESTION = "Is this page a good fit for the tag 'investing'?"
RUNS = 2  # repeat every row to see run-to-run drift


def call_luna(text):
    start = time.perf_counter()
    try:
        d = client.decisions.create(
            model="gpt-6-luna",
            input=text,
            questions=[{"type": "predicate", "name": "fit", "instructions": QUESTION}],
        )
    except Exception as exc:
        return {"status": "error:" + type(exc).__name__, "ms": (time.perf_counter() - start) * 1000}
    ms = (time.perf_counter() - start) * 1000
    a = d.answers[0]
    if a.type == "refusal":
        return {"status": "refusal", "ms": ms}
    usage = getattr(d, "usage", None)
    return {"status": "ok", "ms": ms, "p": a.probability, "tokens": getattr(usage, "input_tokens", 0)}


def pct(values, q):
    s = sorted(values)
    return s[min(len(s) - 1, int(q * len(s)))]


def report(results):
    ok = [r for r in results if r["status"] == "ok"]
    print(f"calls={len(results)} failed_or_refused={len(results) - len(ok)}")
    ms = [r["ms"] for r in ok]
    print(f"p50={pct(ms, 0.5):.0f}ms p95={pct(ms, 0.95):.0f}ms")
    print(f"brier={statistics.fmean((r['p'] - r['label']) ** 2 for r in ok):.4f}")
    print(f"input cost at $0.10/M: ${sum(r['tokens'] for r in ok) / 1e6 * 0.10:.4f}")
    for i in range(10):  # reliability: does 0.9 mean about 90 percent?
        lo, hi = i / 10, (i + 1) / 10
        b = [r for r in ok if lo <= r["p"] < hi or (i == 9 and r["p"] == 1.0)]
        if b:
            print(f"{lo:.1f}-{hi:.1f} n={len(b):3d} predicted={statistics.fmean(r['p'] for r in b):.2f} "
                  f"actual={statistics.fmean(r['label'] for r in b):.2f}")


if __name__ == "__main__":
    rows = [json.loads(line) for line in open(sys.argv[1])]  # {"text": ..., "label": 0 or 1}
    results = [{**call_luna(r["text"]), "label": r["label"]} for r in rows for _ in range(RUNS)]
    report(results)

Point the same loop at Jev or an open model by swapping call_luna for another adapter, keep the labels fixed, and compare the reports side by side.

Honest limits of this verdict

  • Every figure is self-reported. Topfi, agentdev001, and the others are single users with small samples and unpublished methods. We ran no benchmark of our own.
  • Topfi's workload is unusual, and by his own account he had not tested images. A tagging or UI-selection result may not predict your ticket router.
  • The latency conflict is unresolved. We offered hypotheses, not findings.
  • No head-to-head accuracy table exists yet for Luna Decisions against Jev on a shared benchmark panel, as far as we found.
  • Docs gaps: the pages we read give no context limit, image token rates, question or choice limits, or version-pinning policy. The 1M window and the 128-image cap come from a commenter and a plugin README respectively, not from OpenAI.
  • Beta pricing and behavior can change. OpenAI says general availability is expected in the coming weeks.
  • HN timing: the thread was created late on October 6 UTC and was still active on October 7, so some reports may be updated after we read them.

What this means for what you build or pay

If you already ship a decision hop, add Luna as a second provider behind your adapter, mirror a slice of traffic, and compare reliability and dollars per correct decision on your own labels. If your organization needs OpenAI's compliance controls, the day-one numbers do not argue against it, because the price gap is small money at normal volumes. If cost at scale or a pinned model is the constraint, the early evidence favors Jev or open weights, but not conclusively.

Do not threshold on probabilities you have not checked. The surest lesson from day one is Simon Willison's: these models are hard to vibe-check, so the eval is the product decision.

Related on explainx.ai

  • OpenAI Decisions API: GPT-6 Luna at DevDay, plus the public beta update
  • What are decision models? The practitioner category guide
  • Perplexity pplx-decider and its Decisions API
  • How to integrate Jev for agent routing
  • Jev Ultrafast: browser-use with typed decisions
  • llama.cpp adds decision model support
  • Is Jev's speed and cost claim true?
  • Stop using LLMs as classifiers: calibration guide
  • Official: Decisions API docs · OpenAI community announcement · Hacker News thread · llm-openai-decisions on GitHub

Figures come from OpenAI's Decisions API documentation and community announcement, the Hacker News thread "Decisions API is in public beta" (created October 6, 2026 UTC), and the llm-openai-decisions repository, as of October 7, 2026. Commenter numbers are their own reports and were not independently verified. The public beta may change pricing, limits, and behavior. Run your own evals before you ship.

Spotted something out of date? Let us know.

People in this article

  • Simon Willison →Independent open source developer and creator of Datasette
Explore people in AI →
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 3, 2026

llama.cpp Adds Decision Model Support for Five Open Models

On October 2, 2026 the ggml team announced decision-model support in llama-server. You POST state plus typed questions to /v1/systemone and get option probabilities in one forward pass. Five open models ship as GGUF today — and the local runner map just got a third path next to Ollaya and Python/MLX stacks.

Oct 2, 2026

Perplexity pplx-decider: Open Weights Plus a $0.04 Decisions API

On October 1–2, 2026 Perplexity published Apache-2.0 weights for pplx-decider-v1-27b, a ~26B multimodal fine-tune of Qwen3.8-27B, and launched a Decisions API at $0.04 per million input tokens with free output. On Perplexity's 11-benchmark panel the hosted model scores 85.71% overall versus Jev at 84.51% and the Qwen base at 74.76% — and still loses several of the hardest rows to Jev.

Sep 29, 2026

Jeff: Home-Trained Jev-Compatible 0.8B Decision Models

A Hacker News front-page project from firelex ships Qwen3.5 and Gemma fine-tunes that speak Jev's request format locally. The 2B's five-benchmark panel score is 83.1% against Jev's published 83.0% on a different sample; the 0.8B is ~28 ms on an M4 Max. This post is about Jeff specifically — not Kev, OpenJev, or Laya — and what the thread actually argued.