Related — October 6, 2026: Interfaze released an Apache 2.0 model for OCR, speech and extraction with confidence scores. See Interfaze-1-Lite.
Most LLM calls in production are not conversations. They are decisions: is this ticket urgent, does this photo show a defect, should this message be escalated. For those, generating tokens, parsing JSON and paying for output is a roundabout way to get one label. Liquid AI's d1 takes the direct route: it returns a probability for every possible answer, with zero generated output tokens, and as of the announcement covered on October 6, 2026, it now accepts images as well as text.
Liquid's headline claim is aggressive: d1 matches or beats GPT-6.1 Sol on four of six real applications while costing 19 to 200 times less than both GPT-6.1 Sol and Claude Opus 5.5. This post explains what a decision model is, what the numbers do and do not show, and how to test it before moving any workload.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is d1? | A model that returns probabilities over a fixed set of answers, not text. |
| What is new? | Image input, alone or with text, in a single call. |
| Price? | $0.04 per million input tokens, no output tokens billed. |
| Speed? | Text decisions in about 200 to 300 ms, per Liquid. |
| Claimed accuracy? | Matches or beats GPT-6.1 Sol on 4 of 6 tasks; 85 to 97% on four inspection tasks. |
| Open weights? | Not stated. Model size not disclosed. |
| Should I switch? | Only after testing on your own labeled data. |
What a decision model is
We covered the category in our explainer on decision models. The short version: ask a normal LLM "is this email spam?" and it generates tokens like "Yes, this looks like spam because..." that you then parse. A decision model skips generation. It reads the input and returns probabilities for each allowed answer, in one forward pass.
d1 supports three question shapes, per Liquid:
- Yes/no: a probability between 0 and 1.
- Choice: pick one label from several, with a probability for each.
- Score: a position on a scale, weighted by probability.
Three practical benefits follow from that design.
- No output tokens. You are billed for input only, and you cannot be surprised by a long answer.
- No parsing failures. There is no JSON to break, so malformed responses disappear as a failure mode.
- Calibrated confidence. Because you get probabilities, you can set thresholds: auto-approve above 0.95, route to a human between 0.5 and 0.95, and so on.
The category is moving quickly. Perplexity published its own decision API, and llama.cpp added decision model support for five open models. d1 is the commercial offering from Liquid, which has focused on efficient architectures.
What changed: vision
The new version of d1 takes an image, text or both. Liquid says it is the first decision model to extend its text capabilities to visual input. The reported results are in industrial inspection: 85 to 97 percent accuracy on four VisA industrial inspection tasks with no task-specific training, and 97 percent on a factory inspection task where it costs up to 200 times less than GPT-6.1.
Image pricing is simple: 1.5 tokens per 32 by 32 pixel patch, so a 1024 by 1024 image is 1,536 tokens, or about $0.00006 at the listed input price. Even at high volume, that is small. A million inspected images at that size would cost roughly $61 in input tokens by this arithmetic, which is ours and ignores any platform fees.
The headline claim, read carefully
"Matches or beats GPT-6.1 Sol on four of six real tasks, 19 to 200 times cheaper" is a strong sentence. Here is what it is and is not.
What it supports. For well-defined classification tasks with a fixed label set, a small purpose-built model can be competitive with a frontier model and far cheaper. That is plausible and consistent with how specialized models have long beaten generalists on narrow tasks.
What it does not show.
- Which two tasks it lost. Four of six means two where it did not match. Those may be the harder or more ambiguous ones, and the post does not give a full benchmark table.
- How the tasks were chosen. They are applications Liquid selected, such as support-ticket filtering, circuit-board inspection, Tetris and Wordle. Vendor-chosen tasks tend to fit the vendor's strengths.
- How the baselines were prompted. A general LLM given a better prompt, a few examples or constrained output may close some of the gap.
- Cost comparisons. The 19 to 200 times range depends on the task and on how much output the larger models generate. A frontier model asked to reply with one word narrows the gap on output cost, though not on input.
Two of the reported demos are worth reading as demonstrations, not benchmarks: a Tetris agent whose line count rose from 70 to 81 and a Wordle solver that won 12 of 12 games. They show the model can drive a loop of repeated decisions cheaply, which is different from showing it beats a frontier model at open-ended play.
Speed and where it fits in a pipeline
Liquid reports text decisions in 200 to 300 milliseconds. That matters because decisions often sit in the hot path: moderation before a message posts, routing before an agent starts, a gate before an expensive call. A cheap, fast, calibrated classifier can sit in front of a costly model and cut spend.
A common architecture:
input -> d1 (cheap classifier, probabilities)
p(confident, easy) -> act automatically
p(uncertain) -> send to a frontier LLM
p(high risk) -> send to a human
This is the same pattern we describe for agent approvals in our Codex Auto-review guide: put a cheap check up front and reserve expensive attention for the exceptions.
How to test d1 on your own workload
Do not rely on the vendor's six tasks. A one-afternoon evaluation is enough to decide whether to pursue it.
- Pick one decision you already make with an LLM. Ticket routing, moderation, defect detection or lead scoring all work.
- Collect 200 to 500 labeled examples, including the awkward ones. Hold out a portion you never tune against.
- Run your current model and d1 on the same set. Record accuracy, per-class errors, latency and cost per thousand.
- Check calibration. When d1 says 0.9, is it right about 90 percent of the time? Plot or tabulate it. Calibration decides whether thresholds are safe.
- Look at the errors. Which cases does each model miss? If d1 misses the expensive ones, the saving is false.
- Try the cascade. Send only uncertain cases to the bigger model and measure the blended cost and accuracy.
- Check data handling. Confirm where images and text are processed, retention policy and any compliance requirements, especially for inspection images of proprietary products.
Limits to keep in mind
- Closed sets only. If the answer is not one of your labels, d1 cannot say so unless you include an "other" label. Design the label set carefully.
- No explanation. You get a number, not a rationale. For regulated decisions that need reasons, pair it with something that can explain.
- Distribution shift. Zero-shot accuracy on public inspection datasets may not carry over to your parts, lighting or camera angles.
- Undisclosed details. Liquid does not state the model's size or whether weights will be released, so planning for self-hosting is premature.
- Platform availability. Vision is described as coming soon to Vercel and OpenRouter, so the console is the reliable path today.
Where decision models fit among your options
It helps to place d1 next to the alternatives you already have. A frontier LLM gives the most flexible reasoning and explanations, at the highest price and latency. A small open model fine-tuned on your labels can be cheaper still and fully under your control, but you must collect data and maintain it. A classical classifier, such as gradient-boosted trees on features, can be excellent for tabular problems and needs no language understanding. A decision model like d1 sits between them: zero-shot flexibility on text and images, probabilistic output, low price, and no training pipeline on your side.
The choice usually comes down to whether you have labeled data and a stable label set. If you do and the volume is huge, fine-tuning may win on cost. If labels change often or you have little data, a zero-shot decision model is attractive. In both cases, measure calibration, because probabilities that are not trustworthy are worse than none.
What this means for what you build or pay
If a meaningful share of your LLM bill comes from classification, routing or scoring, a decision model deserves a test, because the cost difference is large enough to matter even if the real-world gap is smaller than the claim. A realistic outcome is not "replace GPT-6.1 Sol" but "put d1 in front of it and send it a fraction of the traffic." For the frontier side of that cost picture, see our breakdown of GPT-6.1 Sol pricing and benchmarks.
Related reading
- What are decision models? AI classifiers guide
- Perplexity's decision API
- llama.cpp adds decision model support
- GPT-6.1 Sol launch, pricing and benchmarks
- Codex Auto-review: cheap checks before expensive attention
- Self-improving harnesses and small models (YC panel)
Primary: Liquid AI, "Introducing d1: the most capable decision model, now with vision" (liquid.ai) · MarkTechPost coverage of d1 (September 29, 2026)
Details are accurate as of October 6, 2026 and rely on Liquid AI's own announcement. Benchmark results are vendor-reported on vendor-chosen tasks, and pricing and availability may change.
