Perplexity just put a priced, open-weight model on the hop explainx.ai has been tracking since Jev and OpenAI Decisions API: do not ask a chat model to write a paragraph when the product only needs a closed-set probability.
On October 1, 2026 the company published pplx-decider-v1-27b on Hugging Face under Apache-2.0 — a ~26B multimodal fine-tune of Qwen3.8-27B, about 49 GiB on disk, with a decision readout and an inference.py that speaks choice and noul. The same day it launched a hosted Decisions API. Perplexity's CEO said the API is $0.04 per million input tokens, output free, and that the company intends to cut the input price further. The developer announcement passed roughly 107.8K views.
The score people will paste is real and incomplete. On Perplexity's own 11-benchmark table, pplx-decider-v1-27b is 85.71% overall versus Jev 84.51% and Qwen3.8-27B 74.76%. It wins the panel and several retrieval-ish rows. Jev still wins WinoGrande, BBH, JevBench public hard, TruthfulQA binary, Belebele, and JudgeBench. That split is the story, not the headline average.
TL;DR: What people are asking
| Question | Direct answer |
|---|---|
| What shipped? | Apache-2.0 pplx-decider-v1-27b plus a hosted Decisions API. |
| Price? | $0.04 / million input, free output; CEO said they intend to cut further. |
| Overall score? | 85.71% (pplx) vs 84.51% (Jev) vs 74.76% (Qwen3.8-27B) on 11 benches. |
| Does it beat Jev everywhere? | No. Jev leads WinoGrande, BBH, JevBench hard, TruthfulQA, Belebele, JudgeBench. |
| Local run? | Python 3.12+, CUDA, ~49 GiB weights plus working memory; Decider.predict. |
| Images? | Yes — images=["screenshot.png"]. Config is multimodal. |
| Context? | max_position_embeddings is 262,144 (the ~250k window from the launch posts). |
| Same as OpenAI Decisions API? | Same hop, different vendor. OpenAI: Luna preview, no public price. Perplexity: weights + $0.04/M. |
| System One? | No. Fine-tune of Qwen with a readout, not a non-generative architecture. |
| License? | Apache-2.0 on the Hub card. |
What actually landed on the Hub
The official card is short. The files are not.
perplexity-ai/pplx-decider-v1-27b is tagged license:apache-2.0, base_model:Qwen/Qwen3.8-27B, multimodal, and text-classification. Safetensors metadata lists 26,085,330,160 BF16 parameters. The repo stores eleven weight shards plus readout.safetensors (a ~2.6 MB decision head) and decision_config.json. That config names the training run sft-curated-02, pins the Qwen revision, and stores a fitted temperature of 2.2076. The card tells you the pplx numbers were measured through the Perplexity API, not that your local Decider will reprint them.
config.json is a Qwen3.5 multimodal stack: 64 text layers, a qwen3_5_vision tower, image and video token IDs, and max_position_embeddings: 262144. That is the 250k-class context the launch posts advertised. It is a position-embedding ceiling, not a promise that every 200k-token ticket routes as cleanly as a 2k one.
The repo also vendors AutoJev source under source/src/autojev/ — model.py, types.py, train and evaluate scripts. That is the same family of closed-set fine-tunes explainx.ai already covered with Jeff: a generative Qwen (or Gemma) backbone, letter codes, a temperature, probabilities out. Perplexity's 27B is the first time a frontier search vendor has published that shape at this size with a priced API next to it.
Local inference is not a laptop demo. The card asks for Python 3.12+ and a CUDA GPU with room for approximately 49 GiB of weights plus working memory. The suggested path is uv:
uvx --from huggingface-hub hf download perplexity-ai/pplx-decider-v1-27b inference.py --local-dir .
uv run inference.py
How you call it: choice, noul, screenshot
The product contract is Decider.predict(state, question, images=…), not a chat completion that you later regex.
from inference import Decider
model = Decider.from_pretrained("perplexity-ai/pplx-decider-v1-27b")
result = model.predict(
"My Stripe integration keeps failing. Please help ASAP.",
{
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Charges and refunds",
"technical_support": "Integration errors",
"sales": "Questions about buying a product",
},
},
)
print(result) # Selected choice and calibrated probabilities.
Yes/no is a different question type, not a two-key choice:
result = model.predict(
"My Stripe integration keeps failing. Please help ASAP.",
{"type": "noul", "instructions": "Does this message express urgency?"},
)
Images are a kwarg, not a second model. Pass images=["screenshot.png"] into predict, or uv run inference.py --image screenshot.png. The bundled example asks for a dominant color over {red, green, blue, other}.
The vendored types.py also defines a score question (ordered criteria, a float plus a legend). The public README only documents choice and noul. If you need a Likert-style score, read the source before you assume the hosted API exposes it.
That is the same hop as wiring Jev into an agent for routing and the same hop as OpenAI's Luna-constrained Decisions API: a state, a typed question, probabilities — not a drafted reply.
The 11-benchmark table, without the average washing the losses
Perplexity's launch chart frames a fixed panel (the figure caption uses 7,210 rows, September 2026). The scores below are the Hugging Face markdown table, not the garbled OCR of that chart. Bold is the best score in the row. pplx results are API-measured.
| Benchmark | Jev | Qwen3.8-27B | pplx-decider-v1-27b |
|---|---|---|---|
| WinoGrande | 90.70% | 73.10% | 83.30% |
| FinancialPhraseBank | 76.98% | 75.68% | 84.18% |
| RAGTruth | 77.27% | 61.53% | 88.80% |
| JudgeBench | 78.57% | 68.86% | 78.29% |
| BBH | 94.27% | 72.80% | 82.80% |
| JevBench public hard | 73.27% | 72.28% | 70.30% |
| TabFact | 89.80% | 78.60% | 90.60% |
| ContractNLI | 77.45% | 80.78% | 80.78% |
| Circa | 84.60% | 87.00% | 89.20% |
| Belebele | 95.00% | 93.20% | 94.00% |
| TruthfulQA binary | 92.00% | 82.80% | 85.40% |
| Overall | 84.51% | 74.76% | 85.71% |
pplx-decider wins the overall average plus FinancialPhraseBank, RAGTruth, TabFact, and Circa. It ties the Qwen base on ContractNLI at 80.78% (both beat Jev's 77.45% there).
Jev still wins WinoGrande (90.70 vs 83.30), BBH (94.27 vs 82.80), JevBench public hard (73.27 vs 70.30), TruthfulQA binary (92.00 vs 85.40), Belebele (95.00 vs 94.00), and JudgeBench (78.57 vs 78.29). Those are not rounding errors on WinoGrande, BBH, or TruthfulQA. They are the rows that look most like commonsense, multi-step, and honesty checks.
The Qwen base is not a close third. It sits 10.95 points under pplx overall and loses badly on RAGTruth (61.53 vs 88.80). Fine-tuning moved the needle. It did not invert Jev on Jev's home turf.

Read this the way explainx.ai already tells people to read AI benchmarks: vendor panel, vendor API, one temperature, one month. Useful for "is this in the same quality band." Not useful as a bake-off you can paste into a procurement deck without your own labels.
$0.04 / million input, free output — and why that number matters
The CEO's quote tweet is the rate card we have: $0.04 per million input tokens, output free, with an explicit intent to cut further. A closed-set call that returns a choice label and a probability vector is mostly input. Free output is the correct SKU for this hop.
Put that next to numbers explainx.ai already published:
| Surface | What you pay (as of this writing) | What you get |
|---|---|---|
| Perplexity Decisions API | $0.04 / M input, $0 output | Hosted pplx-decider; Apache-2.0 weights if you leave |
| TypeSafe Jev | $42 / billion input ($0.042 / M), $0 output | System One model; no public 27B weights |
| OpenAI Decisions API | Unpublished (limited preview) | GPT-6 Luna constrained to a finite answer set |
Perplexity and Jev are within a tenth of a cent per million input tokens on the published cards. The fight is not "cheap vs expensive." It is vendor, architecture, and whether you can take the weights home. OpenAI is still the missing price — see the DevDay Decisions API write-up.
If you already route coding traffic the way AT&T described with LiteLLM, this is another cheap hop you can put behind the router: classification, ticket team, "is this urgent," "which tool next." You do not send those turns to a $10 / M writer. You also do not pretend a 27B local decider is free — 49 GiB of BF16 is a real GPU bill.
When that hop starts failing in production, you want traces on the gateway, not another chat log. That is the job LiteLLM Lens is aiming at: grouped failures on the proxy you already run.
Same hop as OpenAI and Jev. Different machine.
explainx.ai's read has not changed since DevDay. Closed-set probabilities are a product category now. Three vendors, three implementations:
OpenAI Decisions API still runs GPT-6 Luna. You define a question and a finite answer list, send text or images, and get a selection in a few hundred milliseconds (vendor claim). Limited preview. No public price. Still an LLM told not to generate.
Jev is the System One bet: it cannot emit free text. Choice, Score, or Noul, calibrated, ~100 ms in TypeSafe's materials. Browser Use already put that into a step loop in Jev Ultrafast. The speed and cost claims are self-tested and should stay caveated.
pplx-decider is a Qwen fine-tune with a readout. noul in the script is the same yes/no primitive Jev named. choice is the same enum hop. Images work. Context is 250k-class. Weights are Apache-2.0. Architecture is closer to Jeff than to Jev. If your compliance story was "this component cannot generate text," Perplexity did not give you that invariant. If your story was "I need a priced API and a fallback I can snapshot_download," this is the first time that sentence is true at 27B from a large vendor.
Do not merge the three into one Terraform module. Map your enum once. Adapter per provider. Fail over on confidence, not on brand.
What people are asking (and the honest limits)
Is 85.71% "better than Jev"? On this panel, overall, yes, by 1.20 points. On WinoGrande, BBH, JevBench public hard, and TruthfulQA binary, no. JudgeBench is a 0.28-point Jev lead. Belebele is a 1.00-point Jev lead. If your production task looks like RAG faithfulness or table fact-checking, the pplx wins are the relevant ones (RAGTruth 88.80 vs Jev 77.27 is the loudest). If it looks like hard commonsense or honesty, Jev's rows are the relevant ones.
Can I treat the Hugging Face checkpoint as the API? The card says pplx scores were measured through the Perplexity API. Local Decider uses readout.safetensors and the saved temperature. That should be close. It is not a guarantee. If you are buying the 85.71% number, buy the API — or re-run the panel yourself.
Will 250k context save me from retrieval? Unlikely. Qwen3.8-27B already advertises a huge window. Perplexity's own Computer write-up has said on-device 27B models start to struggle past ~100k tokens even when the config allows more. Use the window for a long ticket plus a screenshot, not as a reason to skip chunking.
Is AutoJev in the repo a training recipe I can copy? The source tree is there (train.py, evaluate.py, sft_pipeline.py). It is custom code, not a one-click recipe, and the Hub listing is custom-code. Budget time. Jeff already showed a smaller team can land a Jev-shaped fine-tune. A 27B SFT is a different GPU conversation.
Do I need API keys to try the weights? Not for a local load. You need a Hugging Face download path and a GPU. The hosted Decisions API is the $0.04/M meter.
Should I rip Jev out tomorrow? Only if Jev was a single hop with an adapter. Keep generation on a writer model. Swap one enum. Log override rate for a week. That is the same advice we gave when OpenAI previewed Decisions API.
What this means for builders
If you already implemented the contract — state, question, criteria, selection, confidence, human override — add Perplexity as a third provider, not a rewrite. The schema in inference.py is close enough to Jev/OpenAI that the weekend is mapping keys, not redesigning the agent.
If you have not implemented that contract, do it now. Chat-plus-JSON-schema is how this hop silently burns tokens. Structured outputs were the workaround. Decisions APIs are the productization. Perplexity's twist is that the productization comes with weights you can leave with.
Price the GPU honestly. $0.04/M hosted is cheap. 49 GiB local is not "free open source." Route with the same discipline AT&T used on coding models: cheap hop when the label set is closed, expensive hop when you need prose.
Update — October 2, 2026: Category context — What are decision models? places pplx-decider next to Jev, Clef, and LLM-logprob baselines.
Related on explainx.ai
- OpenAI Decisions API: GPT-6 Luna constrained routing
- Jev Ultrafast: browser-use with typed decisions
- How to integrate Jev for agent routing
- What is a System One model?
- Jeff: home-trained Jev-compatible decision models
- AT&T cut AI coding costs 56% with model routing
- LiteLLM Lens: agent trace analysis on the gateway
- TypeSafe AI launches Jev
- Official weights: pplx-decider-v1-27b on Hugging Face
Model card, license, parameter count, Decider.predict examples, and the 11-benchmark table are from perplexity-ai/pplx-decider-v1-27b as of October 2, 2026 (Hub revision 5117a6c, created October 1, 2026). Pricing and the intent to cut input rates further are from Perplexity's CEO announcement; the developer Decisions API post is the ~107.8K-view launch note. pplx scores were measured through the Perplexity API per the card. Verify live docs, rate limits, and your own evals before you ship.
