explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: what people are asking
  • What is Strands Decider 2B?
  • How does the architecture work?
  • How good is it? Accuracy, calibration and latency
  • Why 2B, and what does "open" include?
  • How do I run Strands Decider locally?
  • Strands Decider vs Jev vs pplx-decider vs OpenAI Decisions
  • What developers on Hacker News are saying
  • What are early users reporting on GitHub?
  • Limits the model card tells you to respect
  • What this means for builders
  • Related reading
← Back to blog

explainx / blog

Strands Decider 2B: A Small Open Decision Model You Can Run Locally

Strands Agents, Decision Models, Open Source, AI Agents, On-Device AI

Strands Decider 2B is a 1.9B Apache-2.0 decision model: pointer head, 115 ms median on an RTX 3090, training data included. How it compares to Jev.

Oct 7, 2026·20 min read·Yash Thakker
add explainx.ai
go deep
Strands Decider 2B: A Small Open Decision Model You Can Run Locally

A green robotic gripper choosing one of two items from a tray, an illustration of a decision model picking one option from a short list

Strands Decider 2B is a small, open-source decision model from the Strands team, published October 1, 2026 by Marc Brooker, Mike Chambers and Fabio Nonato de Paula in the official Strands Agents post. It has about 1.9 billion parameters, picks one option from a list or rates something on a scale, returns a confidence with every answer, and cannot write text. The authors report a median of about 115 ms per decision on an RTX 3090. The code and weights are Apache-2.0, and the training data sources and training scripts are public. Digital Today, citing SiliconANGLE, reported it as an AWS release.

That makes it a decision model a builder can retrain on one gaming GPU, with the recipe in the open. It lands after TypeSafe's Jev, Perplexity's 27B decider and OpenAI's Decisions API, and it is built for the same job: the cheap, fast, closed-set choices inside an agent loop. This post covers what it is, how the architecture works, what the benchmarks do and do not prove, how it compares on openness, cost and latency, what early users report, and how to run it locally.

TL;DR: what people are asking

table · 2 cols
QuestionDirect answer
What is it?A decision ("system one") model: pick one of N options, answer yes/no, or rate on an ordered scale, with a confidence per answer.
Who built it?The Strands team, in the strands-labs GitHub org, announced October 1, 2026.
Size and base?About 1.9B parameters: Qwen3.5-2B-Base torso, rank-16 LoRA, pointer head of roughly 1M parameters.
Can it write text?No. The language-model head is removed, so it cannot do chat, coding or summarization.
License?Apache-2.0 for the code and the v19 weights; training data sources, licenses and scripts are in the repo.
How fast?Median 115 ms and p95 299 ms on an RTX 3090; warm median 153 ms for short tasks on an M3 Pro. A CPU takes seconds.
How accurate?v19 scores 0.723 on the JevBench public set (167 of 231). Easy tier is 1.000, hard tier is 0.505.
Which checkpoint?The post pins v19. The repo README now calls v21 the reference (0.762, 176 of 231) and keeps v19 published.
What does it cost?No per-call fee. You pay for hardware you already own. Training takes about 11 hours on one RTX 3090.
Jev compatible?It serves POST /v1/systemone, and the JevBench adapter for Jev-style models ran against it unmodified.
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What is Strands Decider 2B?

The Strands team describes it as one of a new class of "decision models or system one models". Unlike an LLM, which generates arbitrary text, a decision model answers questions such as "Is the string 'turn on the lights' about the coffee machine? Yes or no," or "What language is 'sihamba ngokushesha' in: English, Zulu or Dutch?" It can also score, for example how positive a sentence is between 0 and 1.

Three question types come out of the same machinery, which the repo README names noul (yes/no), choice (one of N) and score (an ordered rubric). If those terms are new, the category is mapped in explainx.ai's guide to decision models and the System One explainer.

The trade is explicit in the announcement. Decision models are faster and more capable at a given size, always return one of the offered options, and run at very low latency. The flip side, in the authors' words, is that generating every output in one parallel pass makes them significantly worse than reasoning models on complex problems, and the lack of text generation rules out coding, chatbots and document summarization.

Two properties are the real pitch for agent builders. Each decision carries a reliability score, which the post says frontier LLM inference APIs do not expose. And asking several questions about the same text is cheap, because the text is read once and each extra question adds only its own tokens, according to the repo's inference docs.

How does the architecture work?

The recipe is deliberately simple. Take a pretrained decoder, here Qwen3.5-2B-Base, discard its language-model head so it cannot generate, and put a small pointer head on top. The torso is adapted with a rank-16 LoRA adapter, and the head runs in fp32.

Strands Decider 2B architecture diagram: a Qwen3.5-2B-Base torso with LoRA r16 feeds hidden states at each option and at the answer position into a pointer head that outputs one logit per option

Figure: the Strands Decider 2B architecture. Source: Strands Agents blog, "Introducing Strands Decider 2B", October 1, 2026 (strandsagents.com).

The input is laid out as a state, a question and a numbered option list, ending at an <answer> marker. The pointer head scores each option by comparing the hidden state at the last token of that option's line with the hidden state at the <answer> position. In the diagram the logit is a scaled dot product between a query projection of the answer state and a key projection of each option state, followed by a softmax. One forward pass gives per-option probabilities, with no decoding loop.

Because the head holds no per-option parameters, the README notes three consequences. Nothing can learn that "the first option is usually right," nothing caps how many options a question can carry, and label sets are defined by the request rather than baked into the weights. That last point is why the same checkpoint can route support tickets in one call and grade an answer for adequacy in the next.

The authors also say this is the second major iteration. The first used a slot head that mapped the final hidden state to a fixed set of slots and performed significantly worse, and the released model is v19 out of many iterations. The same pointer-head idea appears in the open Jev clones, for example Kev, which also pairs LoRA with a pointer head on Qwen3.5.

How good is it? Accuracy, calibration and latency

The team tracks three targets: accuracy, calibration and latency. Accuracy and calibration are measured on the public set of JevBench, an independent benchmark for Jev-class models, with calibration reported as a Brier score. The post's headline is that v19 is third of 33 in the 2B class and first of 30 if you exclude the just-over-2B models. The ranking comes from the JevBench board, and for background on the benchmark itself, see Jev Playground and JevBench.

Scatter chart of JevBench accuracy against Brier score for Strands Decider 2B versions v5 through v19, with an arrow path moving toward higher accuracy and lower Brier score

Figure: accuracy and Brier score across the training trajectory on the JevBench public set (lower Brier is better). Source: Strands Agents blog, October 1, 2026.

The repo's own evaluation pages give the numbers behind that headline, and they add context the announcement leaves out.

table · 3 cols
Measurementv19v21
JevBench public accuracy0.723 (167 of 231)0.762 (176 of 231)
Brier score / expected calibration error0.342 / 0.0520.323 / 0.064
Easy / standard / hard tier1.000 / 0.875 / 0.5051.000 / 0.931 / 0.550
Latency, RTX 3090, median / p95115 ms / 299 mssame architecture and size as v19
Latency, M3 Pro, warm median, short tasks / all tasks153 ms / 234 mssame architecture and size as v19

Three honest readings follow from that table and the surrounding pages.

First, "3rd of 33 in the 2B class" is a class ranking. The same repo page places v19 50th overall of 89 ranked systems on the JevBench v1.4.2 board from September 25, 2026, and says the top is reasoning models plus Jev rebuilds on larger torsos. The 100% score applies to the easy tier, which the repo says most systems saturate. The hard tier, at 0.505, is where a 2B model falls well short of the leaders.

Second, the board has moved. When explainx.ai checked the JevBench page on October 7, it showed a v1.6.1 board, and we could not find Strands Decider 2B on it. Treat the rank as a September 25 snapshot from the maintainers' repo, not a live standing.

Third, noise matters. The maintainers write that six retrains of one recipe varied by a standard deviation of about 3.2 tasks, so differences under roughly 10 tasks between two single runs are unresolved. That is why v21's gain over v19 should be read as a hint, not a result.

What does the confidence score buy you?

The README says that on short classification tasks the model has never seen, answers at a confidence of 0.9 or more are right about 95% of the time, and that below that you should confirm or ask a person. The model card adds a caveat: calibration is one temperature per question primitive, fitted only on held-out short classification, so measure it on your own traffic before trusting a threshold.

Scatter plot of Strands Decider 2B v18 decision latency in milliseconds against input tokens on an RTX 3090, flat near 100 ms for short prompts and rising to about 300 ms near 3,000 tokens

Figure: decision latency against task size on a local RTX 3090 for v18. Source: Strands Agents blog, October 1, 2026.

What does the latency chart actually show?

Reading the chart, latency sits near a 100 ms floor for prompts up to roughly a thousand tokens, then climbs to about 250 to 300 ms at 2,500 to 3,000 tokens. The chart caption reports 230 requests through the local server with the HTTP round trip included, a p50 of 106 ms, a p95 of 296 ms, and a first request of 7.7 seconds that was left off the plot as server warm-up.

The Mac number needs a footnote. The repo's benchmark shows the 153 ms median applies to repeated requests under 300 tokens on Apple's GPU, because the backend compiles per input length. The first request at a new length took a median of 310 ms for those short tasks, and across all 231 tasks the first-request p50 and p95 were 622 ms and 3,862 ms against 234 ms and 2,628 ms when repeated.

Why 2B, and what does "open" include?

The authors give two reasons for 2 billion parameters. It is small enough that you can serve, and even retrain, the model on hardware you already own, and it looks like a sweet spot between experimentation and meaningful work. The repo adds the numbers: about 11 hours to run the full training recipe on one RTX 3090, or about 1 hour 10 minutes on eight H100s with the fast flag, using training/recipe.sh or the AWS runner.

"Open" here is broader than weights. The v19 model card lists the public datasets in the training mix, including GLUE, PAWS, AG News, ContractNLI and MuSiQue, plus synthetic rows generated and checked by open-weight models and output distributions from a frozen Qwen3.5-4B teacher. The code repo has data/sources.md listing each source with its license, and a preregistered research log: since v9 every run states predictions and failure conditions before training, and most runs missed their bar, v20 among them.

That is a lot of disclosure for a launch. It also means the license question has a second half, which an early GitHub issue already raised (see below): the base model is Alibaba's Qwen, and some data-generation steps used Qwen too.

How do I run Strands Decider locally?

The easiest path is the strands-decider CLI. The commands below come from the official post and the repo README; explainx.ai did not run them for this article, so treat the printed output as the authors' example.

bash
pip install strands-decider

strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
  --state "Help! My payouts have been failing for 3 days! " \
  --choice "Which team should handle this?=billing,sales,retail"

The post shows output like the following, with the winning option, a confidence, and a probability per option:

text
choice_0 -> billing (confidence 0.768)
  billing   0.845
  retail    0.091
  sales     0.064

Notice that confidence (0.768) and the top option's probability (0.845) are different numbers, so log both. The CLI also accepts --noul "Does this convey urgency?" for yes/no and --score "How frustrated is the writer?=calm,frustrated,depressed" for an ordered scale, and you can combine several questions in one call so the state is loaded once.

Serving it as an endpoint

strands-decider serve starts a local HTTP server on the same /v1/systemone path that llama.cpp uses for its decision-model support. The README's request looks like this:

bash
strands-decider serve StrandsAgents/strands-decider-2B-hobson-v21 --port 8000

curl -s localhost:8000/v1/systemone \
  -H 'content-type: application/json' \
  -d '{
    "state": "Help! My payouts have been failing for 3 days!",
    "questions": {
      "is_urgent": {"type": "noul", "instructions": "Does this convey urgency?"}
    }
  }'

The response returns the answer, token usage and a latency_ms field. On an Apple-silicon Mac, --device mlx runs the model through MLX, which the README says is 1.4 to 1.6x faster than MPS, but it needs the mlx extra and, until the next release, an install from a clone. Because the base Qwen3.5-2B is natively multimodal, a --vision flag lets requests carry images, and the README says the published checkpoints answer over them without image training.

Using it as a gate inside a Strands agent

The repo's examples/strands/ folder shows the pattern the authors care about. An agent with a get_weather tool is told to be eager, so when a user asks "What's the weather?" with no city, it guesses one. Before the tool call runs, the decider reads the conversation and the proposed call and answers two yes/no questions: are the argument values grounded in something the user said, and is it premature to call this tool before clarifying.

The seam is Strands' intervention system. You write an InterventionHandler with a before_tool_call method and pass it to the agent. It returns a typed action: Proceed to let the call run, Deny to block it, Confirm to stop and ask a person, or Guide to hand the model its turn back with feedback. The example uses Proceed and Guide, and the agent asks which city instead of reporting weather nobody asked for.

python
QUESTIONS = {
    "args_grounded": Decider.noul(
        "Are the tool's argument values grounded in facts the user actually provided?",
        {"true": "every argument value traces back to something the user said",
         "false": "an argument value was guessed or invented, not stated by the user"},
    ),
    "premature": Decider.noul(
        "Is it premature to call this tool now, before clarifying with the user?",
        {"true": "the assistant should ask a clarifying question before calling the tool",
         "false": "there is nothing left to clarify; calling now is appropriate"},
    ),
}

The example needs AWS credentials with Bedrock access because the agent's own LLM runs on Amazon Bedrock, and the README tells you to serve the decider on port 8099. The authors are explicit that the questions, thresholds and policy were picked by hand and that this is an illustration, not a recommendation. The same shape applies to the hybrid-agent idea in the post: let an LLM make the hardest calls and a decider handle the rote ones. For where to place such checks in a pipeline, see cheap verification checkpoints with Jev.

Strands Decider vs Jev vs pplx-decider vs OpenAI Decisions

All four target the same hop: a state, a typed question, probabilities out. They differ on openness, price and where the compute runs. The table uses figures explainx.ai has already documented in its own coverage, linked in the table, and the vendors' numbers were not measured under one harness.

table · 5 cols
Strands Decider 2BJev (TypeSafe)pplx-decider-v1-27bOpenAI Decisions API
AccessLocal weights, Apache-2.0, plus training data and scriptsHosted API (launch coverage)Apache-2.0 weights plus hosted API (coverage)Closed, public beta (coverage)
SizeAbout 1.9BNot stated in our coverageAbout 26B, roughly 49 GiB of GPU memoryGPT-6 Luna, an LLM constrained to your options
PriceNo per-call fee, your hardware0.042 dollars per million input tokens, output free0.04 dollars per million input at launch, output free0.10 dollars per million input, output free
Latency claimMedian 115 ms on an RTX 3090, 153 ms warm on an M3 Pro70 ms to 500 ms end to end, self-reported (fact-check)Not published in our coverage"10x faster than the Responses API"
Can it generate text?NoNo, by designFine-tune of a generative Qwen with a decision readout, per our coverageIt is an LLM, constrained
ImagesOptional, via a vision flagText only at launchYesYes, inline base64

What does the money look like?

At a typical 300 input tokens per call, one million decisions is 300 million input tokens. At the listed rates that is about 12.60 dollars on Jev, 12 dollars on pplx-decider at its launch price, and 30 dollars on OpenAI Decisions. Strands Decider has no such meter, but it is not free: you supply the GPU or Mac, and there is a throughput ceiling.

At the published 115 ms median, one sequential stream handles about nine decisions a second, which is roughly 32 hours for a million single-question calls. Batching several questions over one shared state helps, but the server concurrency issue described below means you cannot simply fan requests out today. For a burst workload, a hosted API still wins on elasticity. For a steady agent loop on hardware you already own, a local 2B model wins on marginal cost, privacy and the ability to fine-tune.

Where does Strands Decider lose?

Accuracy, against the larger systems. The repo itself says the top of the JevBench board is reasoning models and Jev rebuilds on bigger torsos, and where Jev actually fails is a reminder that even the leaders have gaps. If your decisions are hard, multi-step or involve long documents, a 2B model scoring 0.505 on the hard tier is the wrong default. For other small local options, compare Julia 1 on a CPU, Laya on MLX and the Ollaya runtime.

What developers on Hacker News are saying

The announcement was submitted to Hacker News on October 7, 2026 and had 70 points and a handful of comments when we checked, so this is early and light feedback. It is one thread with mostly positive reactions and no head-to-head benchmark reports yet.

  • A working on-device use. One commenter, miguelspizza, said they run it on device in a Chrome extension to filter email, and called it the right mix of size, capability and speed for ad hoc bulk classification. They pointed to a community WebGPU build, which its model page describes as an int4 browser port with the v19 LoRA merged. explainx.ai did not run it.
  • A live browser demo. In reply, babelfish shared a demo link for that browser build, which suggests the model is small enough to run client side.
  • Praise for the writing, and a request for ideas. keyle said the post was readable by non-specialists and asked others to suggest use cases, and no concrete ideas had been posted at the time we checked.
  • A Mac mini use case. davvie thought it could handle smaller automations on a Mac mini.
  • The open CPU question. teruakohatu asked how well it would run on a CPU. No reply had answered it when we looked, but the repo does: on an M3 Pro's CPU in fp32, a 256-token single question took about 3.3 seconds against about 0.3 seconds on the GPU backend, and peak memory was 8.3 GB against 5.4 GB. The post's "tens of milliseconds" is a GPU or Apple-GPU claim, not a CPU one.

Nobody in the thread reported accuracy numbers or compared it with Jev. That gap is the main reason to run your own labeled set.

What are early users reporting on GitHub?

The repo's issue tracker had 39 open items (issues and pull requests combined, per the GitHub API) on October 7, six days after launch. A few matter for anyone planning a deployment.

  • Concurrency can return wrong answers. Issue 17, filed by davetbo-amzn, reports that concurrent requests to serve can read each other's option offsets and return silently wrong probabilities with HTTP 200, or a 422. ghchinoy independently reproduced it on CUDA: with 8 concurrent clients one choice moved by 0.82 probability and flipped, while wrapping the engine call in a lock made results match serial runs and was also faster, at 479 ms against 921 ms p50 server time. It was still open when we checked.
  • A Mac crash on overlapping requests. Issue 9, from yeshwanthlm, reports that serve crashes on Apple silicon when two requests overlap. A team member, nonatofabio, reproduced it on an M4 Pro and identified a second race on the shared option offsets that is not specific to Apple silicon.
  • A 38-second cold start. Issue 22, from ghchinoy, measured the first request after a Cloud Run scale-from-zero at 38.4 seconds end to end, 20.9 seconds of it server time, against about 34 ms for later requests. The proposal is to warm the engine before the health check reports ready.
  • A request for a non-Chinese base model. Issue 2, from davetbo-amzn, says some government-cloud customers would prefer a version based on a non-Chinese open model such as Gemma 4. A team member, nonatofabio, replied that Gemma should be doable and asked whether data lineage would also need to avoid non-US models, since some data-generation steps used Qwen.
  • Language coverage is unknown. Issue 8, from jparadaa, asks whether Spanish is supported or evaluated, because the card lists no languages. No answer had been posted.
  • Unofficial forks are already arriving. Several issues submit locally benchmarked candidates, including a 4B variant claiming a top spot on the JevBench text board, all labeled unofficial by their authors. Treat those claims as unverified.

Limits the model card tells you to respect

The v19 model card is blunt, and it is worth reading before you wire the model into anything that matters.

table · 2 cols
LimitationWhat it means in practice
Questions are read less than documentsWith the state and options fixed, a reworded question often gets the same answer. Phrase it so the obvious reading is the intended one.
Long, multi-step documents are the weak spotThe hard tier is 0.505. Do not use it for multi-hop policy checks without your own eval.
score and noul transfer poorly to unlike rubricsTrain or tune on a rubric that resembles yours.
Calibration is one temperature per primitiveThe confidence bands are validated only on short classification. Measure on your traffic.
Trained on public datasetsIt inherits their domains and label noise.
Window sizev19's figures use a 3072-token window; at the saved 4096 window it scores 168 rather than 167.

What this means for builders

Start by picking one rote decision in your agent that currently costs an LLM call: tool selection, ticket routing, "is this request urgent," a grounding check before a tool runs. Write 100 labeled examples from real traffic, run both v19 and v21, and compare accuracy and the calibration of the confidence field on your own data, not on JevBench.

Then set the policy. Branch on confidence, send anything under your threshold to a bigger model or a person, and log the override rate for a week. The README's 0.9 figure is a starting point from short classification tasks, not a guarantee for your domain.

Finally, respect the operational rough edges. Serialize requests to serve until the concurrency issue is fixed, warm the engine before taking traffic, and plan for a GPU or Apple silicon rather than a CPU. If you want a hosted path with a price list, the other three options in the table are the place to look, and the Jev routing guide shows what that integration looks like. The Strands post says its team is working on integration libraries, so watch the repo.

Details, benchmarks, issue counts and repository state are accurate as of October 7, 2026. explainx.ai did not run the model or the commands in this article; figures come from the Strands Agents post, the strands-labs repository, the Hugging Face model card, and the Hacker News and GitHub threads cited. Hacker News and GitHub comments are individual users' reports, not verified results.

Related reading

  • What are decision models? The practitioner category guide
  • What is a System One model?
  • Perplexity pplx-decider: open weights plus a Decisions API
  • OpenAI Decisions API: GPT-6 Luna constrained routing
  • llama.cpp adds decision model support for five open models
  • Kev: the open-source Jev clone family
  • Jev Ultrafast: Browser Use puts Jev in the browser agent loop
  • How to wire Jev into your agent pipeline for routing
  • Official sources: Introducing Strands Decider 2B, strands-labs/strands-decider on GitHub, strands-decider-2B-hobson-v19 on Hugging Face, Introducing Strands Labs on the AWS Open Source Blog
Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 27, 2026

Julia 1: A 144M Decision Model You Can Run on a CPU

Supersonic Labs released Julia 1, a 144.3 million parameter encoder you can run on a CPU or in a WebGPU browser under Apache 2.0. The published edge over a Jev reference is 0.45 points on one suite, and Banking77 falls to 64%.

Aug 5, 2026

LFM2.5-2.6B: Liquid AI's Biggest On-Device Agent Model Yet

LFM2.5-2.6B is Liquid AI's flagship on-device agent model — 2.6B parameters, 34 trillion training tokens, and benchmark scores that beat Gemma-4-E4B and match Qwen3.5-9B on tool use, while running under 2.5GB of memory on a phone. explainx.ai covers the numbers and where it fits next to Liquid's smaller LFM2.5-230M.

Oct 6, 2026

Octop: Tencent Cloud's Open-Source Multi-User AI Workspace

Tencent Cloud open sourced Octop, a self-hosted AI assistant platform where each person on your machine gets a login, their own agents and an isolated workspace. It hands work to coding agents and lives in chat apps. Here is how it works, how to start, and what to check before you trust it.