Most production AI work is not chat. It is reading an invoice, transcribing a call, pulling fields into JSON, tagging a ticket. Those jobs need the same answer every time and a way to check it, which is not what a creative general model is optimized for. Interfaze, a Y Combinator 2026 company, built a model for exactly that and released it as open weights on October 5, 2026: interfaze-1-lite, under Apache 2.0.
This post covers what the model does, the architecture claim, the price and hardware, how the reported benchmarks should be read, and a short evaluation plan. We have not run the model ourselves, so the performance statements are the vendor's.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | An open-weight hybrid model for OCR, speech-to-text, extraction, classification and detection. |
| License? | Apache 2.0. |
| Inputs? | Text, images, audio and files such as PDFs and Word documents. |
| Outputs? | Text or JSON, plus confidence scores and bounding boxes. |
| Hardware? | One 80 GB GPU for the full model (compute capability 8.9 or newer). |
| API price? | $0.85 per million input tokens, $1.50 per million output. |
| Context? | 128K tokens in, up to 32K out. |
| Parameter count? | Not stated in the coverage we reviewed. |
What "deterministic work" means here
Interfaze draws a line between creative tasks and deterministic tasks, the backend jobs where a wrong answer is a bug and a right answer should be reproducible. Examples it lists: document OCR, speech-to-text, speaker diarization, structured extraction, classification, object detection, GUI detection, translation, forecasting and guardrails.
The design goal is not to be the smartest model. It is to be accurate, checkable and cheap on tasks where you know what a correct answer looks like. Outputs include confidence scores and location metadata such as bounding boxes, so a downstream system can decide what to trust. A typical flow: extract transactions from a scanned statement into a JSON schema, return the OCR text, the position of each field on the page and a confidence for each value, then send anything below a threshold to a human.
This is the same philosophy behind decision models, which return calibrated probabilities instead of prose, as in Liquid AI's d1 and our decision model guide. Interfaze applies it to perception and extraction rather than classification alone.
Mixture of Architecture
The architecture claim is the part that differs from mainstream models. Interfaze says it fuses task-specific deep networks (CNNs and DNNs) directly into a transformer decoder through a shared embedding space, and calls the result Mixture of Architecture, or MoA. A reasoning core is paired with specialists for tasks such as OCR and speech, instead of asking one general network to do everything.
The intuition is old and sensible: a convolutional network is very good at reading pixels, an audio model is very good at speech, and a language model is good at reasoning over the results. Fusing them avoids the pipeline of separate models glued together with glue code, while keeping specialized components where they help. Whether the fusion beats a well-built pipeline in practice is an empirical question. The company's earlier research write-up argues for task-specific small models, and this release is the productized version.
Price, limits and hardware
Reported figures:
| Item | Value |
|---|---|
| API input price | $0.85 per million tokens |
| API output price | $1.50 per million tokens |
| Context window | 128,000 tokens |
| Maximum output | 32,000 tokens |
| Self-hosting | Full model on a single 80 GB GPU (H100 class) |
| Minimum compute capability | 8.9 or newer |
| License | Apache 2.0 |
An 80 GB requirement puts it in reach of a single data-center GPU, not a gaming card. The weights being open means you can run it inside your own perimeter, which matters for documents containing personal or financial data. For the hosted alternative, compare it with document APIs such as Mistral OCR 4 with bounding boxes and open options such as Baidu's Unlimited OCR.
Reading the benchmark claims
The coverage reports that interfaze-1-lite leads on several measures: MMMU-Pro, RefCOCO, structured output, olmOCR and VoxPopuli speech tasks. One concrete number: 83.8% on olmOCR versus 83.5% for a Claude Sonnet 5 baseline, a difference of 0.3 points, which is inside normal run-to-run variation. It is also weaker on OCRBench V2, text-to-SQL and multilingual question answering compared with the company's larger interfaze-1 model.
Three cautions apply to any vendor benchmark table.
- Small margins are not wins. A 0.3-point lead says "comparable," not "better."
- Benchmarks are not your documents. Receipts in your language, with your scanner artifacts and your layout quirks, are the real test.
- The comparison set is chosen by the vendor. Check which models were left out.
The honest takeaway is that a small, specialized, open model is reported to be competitive with large general models on perception and extraction tasks at lower cost, which is plausible and worth testing, not proven.
How to evaluate it on your own documents
- Collect 100 to 300 real samples, including the ugly ones: skewed scans, handwriting, low-contrast photos, multi-column PDFs and noisy audio.
- Define ground truth in your target JSON schema. Decide how to score partial matches.
- Run your current pipeline and interfaze-1-lite on the same set. Record field-level accuracy, not just overall.
- Check confidence calibration. When it reports 0.95 confidence, is it right about 95 percent of the time? Calibration decides whether automatic routing is safe.
- Test the boxes. Do the bounding boxes actually land on the source text? Wrong boxes undermine auditability.
- Measure cost and latency at your volume, including the GPU cost if self-hosting.
- Check failure modes. What happens with a blank page, a rotated page or an unsupported language?
- Pin and version. For repeatability, fix decoding settings and record the exact model version.
Where this fits in a stack
A reasonable architecture uses a model like this as the first pass and a larger model as the exception handler.
document / audio -> interfaze-1-lite (extract + confidence + boxes)
high confidence -> write to database
low confidence -> larger LLM or human review (with the box highlighted)
That pattern keeps costs low and makes errors reviewable, and it mirrors the cheap-check-first approach we describe in our Codex Auto-review guide. It also pairs naturally with agents that need to read forms and receipts as part of a workflow, where a deterministic reader reduces the chance that a hallucinated number flows into an action.
Limits and open questions
- Size and training data are not disclosed in the coverage we reviewed, so licensing of the training data and exact capacity are unknown.
- Multilingual quality is flagged as weaker on some tests, which matters if you process many languages.
- "Lite" implies a bigger sibling. The larger interfaze-1 may do better on harder tasks, with its own price and access terms.
- Open weights need operations. Running an 80 GB model reliably means serving, monitoring and updates, which may cost more than the API at low volume.
- Security. Any model that ingests untrusted files can be targeted by malicious documents, so sandbox the pipeline and validate outputs.
Cost sanity check
It helps to estimate cost before you commit. At $0.85 per million input tokens and $1.50 per million output tokens, a document that becomes 3,000 input tokens and returns 500 tokens of JSON costs about $0.0033, roughly a third of a cent. A batch of 100,000 such documents costs about $330 in API fees by this arithmetic, which ignores any minimums and platform fees. Compare that with your current per-page cost and with the engineering cost of maintaining a multi-model pipeline. If you self-host on a rented 80 GB GPU, the break-even depends on utilization: a GPU that sits idle half the day costs the same as one that is busy.
Privacy and compliance notes
Documents and recordings often contain personal data, which makes the self-hosting option more than a convenience. Running Apache 2.0 weights inside your own environment keeps files from leaving your control, which simplifies data-processing agreements and some regulatory requirements. Keep logs minimal, because extracted text and audio transcripts can themselves contain sensitive information, and apply retention limits to confidence and bounding-box metadata as well.
What this means for what you build or pay
If a large share of your LLM spend goes to reading documents or transcribing audio, a specialized open model is worth a bake-off, because the price gap with frontier APIs can be large and the confidence and box outputs make human review practical. If your volume is small, the API at $0.85 and $1.50 per million tokens is the easier path. If you handle sensitive documents, the Apache 2.0 weights give you an on-premises option, with the operational work that implies.
Related reading
- Liquid AI d1: a decision model with vision
- What are decision models? Classifier guide
- Mistral OCR 4 with bounding boxes
- Baidu Unlimited OCR: long-horizon parsing
- Perplexity's decision API
- Security-One 27B: an open decision model for security triage
- Choosing open-weight vs closed models
Primary: Interfaze announcement and model page (interfaze.ai) · RuntimeWire coverage of interfaze-1-lite (October 5, 2026)
Details are accurate as of October 6, 2026 and rely on vendor and press reports. We have not run the model. Pricing, benchmarks and limits may change.
