explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • What OpenAI actually announced
  • What the demo is actually deciding
  • How you would think about the call (not an official SDK)
  • How fast is "a few hundred milliseconds"?
  • Decisions API vs Jev: same job, different machine
  • What people are asking
  • Honest limits (as of September 30, 2026)
  • What this means for builders
  • Related on explainx.ai
← Back to blog

explainx / blog

OpenAI Decisions API: GPT-6 Luna Constrained Routing at DevDay

OpenAI, DevDay, GPT-6 Luna, AI Agents, APIs

OpenAI Decisions API (DevDay 2026) uses GPT-6 Luna for constrained, sub-second choices. Preview only; pricing unpublished. Not a System One model.

Sep 30, 2026·12 min read·Yash Thakker
add explainx.ai
go deep
OpenAI Decisions API: GPT-6 Luna Constrained Routing at DevDay

OpenAI Developers (@OpenAIDevs) used DevDay 2026 to ship a product that looks, at first glance, like the company finally built a first-party answer to the routing layer builders have been bolting on themselves. The Decisions API is powered by GPT-6 Luna. You define a question and a finite list of answers, send text or images as context, and Luna returns a selection — classify this, route that, pick the agent's next action. Official copy: "Give your app real-time decision-making." The follow-up post is more concrete: a support ticket plus the teams it could go to, and the API picks one.

This is not a second recap of the whole keynote. explainx.ai already covered the day's stack in the DevDay 2026 announcement roundup, including GPT-6.1 Sol pricing and Dots always-on agents. This post is only Decisions API: what OpenAI actually said, what is still unpublished, and why calling it "OpenAI's Jev" is a useful job description and a bad architecture claim.

XSource postOpen on X ↗

TL;DR

table · 2 cols
QuestionDirect answer
What shipped?Decisions API: constrained questions with finite answers, powered by GPT-6 Luna, for classification, routing, and next-action choice.
Who can call it?Limited preview for selected API customers. OpenAI has said a broader release is planned in the coming days.
Pricing?Not published. Do not invent a rate.
Latency?Tuned for under a few hundred milliseconds end to end, including visual context (Tibo). Community chatter: about 10x a standard model call.
Inputs?Text or images as context, plus the question and the allowed answers.
Is it Jev?Same product job (fast constrained choice). Different stack: Luna is still an LLM; Jev is a System One model that cannot emit free text.
Docs / SDK?No public request schema as of September 30, 2026. Treat any code below as a shape, not a copy-paste client.
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What OpenAI actually announced

The official developer account framed the product in one sentence: give an app real-time decision-making with Decisions API, powered by GPT-6 Luna. Define questions and possible answers to classify content, route requests, or choose an agent's next action. Available in limited preview.

The same thread added the context contract. You send text or images. The worked example is a support request plus the teams that could own it; the API returns a selection. Preview is for selected API customers, not every org on the platform.

OpenAI's recap line is the cleanest product definition we have: "Decisions API enables real-time decision-making by focusing Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers." That sentence is doing two jobs. First, it tells you the output space is closed — you are not asking Luna to write a reply. Second, it admits the intelligence is still Luna. The product is a harness around a small model, not a new model family.

Tibo, commenting from OpenAI's side, called it "lightning fast constrained decision making powered by Luna," with visual inputs, "tuned to be able to make decisions in less than a few hundreds of milliseconds end to end." That is the latency claim you will see quoted everywhere. It is a vendor tuning target, not a published p50/p99 table.

What the demo is actually deciding

Community clips after the keynote showed a nonstop-flights finder: a bounded question ("which itinerary / which hop is allowed") over a closed option set, answered fast enough to feel like a UI control rather than a chat turn. That is the point of the product. The interesting work is not "can Luna write a paragraph about flights." The interesting work is "can Luna pick from a list you already enumerated, with enough confidence that you will act on it without a human in the loop."

That is the same job we already document for Jev in how to integrate Jev for agent routing and in Browser Use's Jev Ultrafast browser loop: most agent steps are a Choice, not a novel. The difference is who sits under the Choice.

If you already ship structured outputs and JSON schemas, Decisions API is OpenAI's attempt to productize that pattern as a dedicated surface — fewer tokens of preamble, a first-class question/answer schema, and a latency target that ordinary Chat Completions rarely hit when you also ask the model to think.

Constrained JSON-style choice diagram reused as a stand-in for Decisions API's finite answer set

How you would think about the call (not an official SDK)

OpenAI has not published a public OpenAPI spec for Decisions API as of this writing. The sketch below is a mental model from the official wording — question, finite answers, text or image context, selection out — not a guaranteed field list.

ts
// Illustrative shape only. Official request/response fields are unpublished.

type DecisionRequest = {
  question: string;
  answers: string[]; // finite, pre-defined
  context: {
    text?: string;
    imageUrls?: string[];
  };
};

// Official support example: ticket + teams it could go to → one selection
const supportTriage: DecisionRequest = {
  question: "Which team should own this request?",
  answers: ["billing", "engineering", "trust-and-safety", "general"],
  context: {
    text: "My invoice charged twice after the GPT-6.1 Sol upgrade.",
  },
};

// Visual input: screenshot + allowed next agent actions
const agentNextStep: DecisionRequest = {
  question: "What should the agent do next?",
  answers: ["click_confirm", "ask_user", "escalate", "abort"],
  context: {
    text: "Checkout page after applying a coupon.",
    imageUrls: ["https://example.invalid/checkout.png"],
  },
};

Until docs exist, do not bake this object into production. Use it as a spec for your own wrapper: if OpenAI's preview client arrives with different names (options instead of answers, image instead of imageUrls), your domain types stay the same.

What you should already be logging, even in preview:

  • The exact question string (prompt drift is still prompt drift)
  • The closed answer list, in order
  • Which selection came back, and any confidence field if the preview exposes one
  • End-to-end latency from your edge, not just model time
  • Whether a human overrode the selection

That last line is the product test. If operators override 30% of routes, you do not have a decision API. You have a suggestion API with extra latency theater.

How fast is "a few hundred milliseconds"?

Tibo's bound is less than a few hundreds of milliseconds end to end, including the visual path. That is the interesting clause. Screenshot in, selection out, still under a third of a second if the tuning holds. Standard Luna chat calls are not advertised that way; this surface exists because OpenAI is willing to specialize the decode (or skip most of it) when the answer set is tiny.

Community reactions after the demo clustered around ~10x faster than a standard model call. Treat that as a relative impression from people watching a stage demo, not a reproducible benchmark. We do not have OpenAI's batch size, region, image resolution, or answer-set cardinality for that number.

Compare that, honestly, to the Jev conversation. TypeSafe's own marketing for Jev talks about roughly 100 ms class decisions and large multiples versus "a full LLM call." Those numbers are also vendor-sourced; explainx.ai has already been skeptical of the biggest multiples in our Jev coverage. The useful comparison is not "10x vs 200x." The useful comparison is: both products are betting that routing should not wait on a generative decode.

If your current stack is chat.completions plus response_format: json_schema with four enums, you already know the failure mode. The model still samples. It still sometimes emits an extra key. It still burns a thinking preamble you did not ask for. Decisions API is OpenAI saying that failure mode is now a first-class product to kill — by keeping Luna, but pinning it to the enum.

Decisions API vs Jev: same job, different machine

Tanishq's reaction after the announcement was the line everyone screenshotted: "rofl OpenAI has built a Jev competitor." That is a fair product joke. It is a bad research paper.

Jev, from TypeSafe AI, is a System One model: it does not generate free-form text. It evaluates a state against typed questions and returns a Choice, a Score, or a boolean-style Noul with a calibrated confidence. The output vocabulary is closed by construction. There is no hidden tokenizer loop waiting to write a paragraph if you forget a stop sequence.

Decisions API still uses an LLM — GPT-6 Luna — constrained to a choice set, with a selection and (in OpenAI's framing) a confidence signal. Constrained decoding, a dedicated endpoint, and a latency budget can make that feel like Jev in a support-router demo. Underneath, Luna remains a generative model that has been told not to generate. Those are different reliability stories.

table · 3 cols
DimensionOpenAI Decisions APIJev (TypeSafe AI)
What it isAPI surface on GPT-6 LunaSystem One model (non-generative)
Can it write prose?Luna can, in other products; this surface asks it not toNo — it cannot generate free text
InputText or images + question + finite answersState + typed questions (product-specific)
Latency claimSub-few-hundred-ms e2e (Tibo); ~10x vs standard calls (community)~100 ms class in TypeSafe's own materials
AvailabilityLimited preview, selected API customersSeparate vendor, own integrations
PricingNot publishedPublished on TypeSafe's side (not linked here)
Closest explainx.ai guideThis postJev routing integration

Name Jev in prose when you compare. Do not treat a constrained Luna call as proof OpenAI trained a System One model. If OpenAI later ships a non-generative decision model, that will be a different announcement. This one is not it.

The integration implication is practical. If you already wired Jev into a LangChain TypeSafeClassifier or Vercel experimental_evaluate path, Decisions API is a second provider for the same hop, not a drop-in replacement of Jev's types. You will still need an adapter: map your closed enum to OpenAI's answer list, map the selection back, and decide what happens when confidence is low. If you never adopted Jev, Decisions API is the path of least resistance if you are already on OpenAI auth, billing, and image understanding — once you are off the preview waitlist.

Browser agents sit in the middle. Jev Ultrafast scores operations and DOM targets so a generative model only writes when the step needs language. A Luna Decisions call on a screenshot is the OpenAI-shaped version of that idea: vision in, next action out. It will not automatically inherit Jev Ultrafast's "never execute a hallucinated selector" rule. That rule lives in the harness, not in the model brand.

What people are asking

Do I need a new API key?

Unknown. OpenAI has not said whether Decisions API is a new product SKU, a header on existing Chat Completions, or a preview flag on the same project. Selected API customers implies some allow-list. If you are not on it, you cannot A/B this in production today.

When does the broad release land?

OpenAI said a broader release is planned in the coming days. That is a calendar hint, not a GA date. Preview semantics can change on the day docs go public: rate limits, image size caps, maximum number of answers, whether confidence is guaranteed.

Is Luna the wrong model for this?

Luna is the cheap, fast GPT-6 SKU — 50% cheaper than the prior generation on the standard API. Using it here is consistent: routing should not spend Astra tokens. The tradeoff is the same one evaluators already found on Luna coding benches: cheaper does not always mean smarter. A mis-route to Trust & Safety vs Billing is a different error than a slightly worse SWE score. Measure your confusion matrix, not OpenAI's stage demo.

How is this different from JSON mode?

JSON mode and strict schemas already let you force an enum. Decisions API is OpenAI productizing the intent: one question, finite answers, real-time, multimodal context, latency as a feature. If the preview is just Chat Completions with a nicer wrapper, the 10x claim will not survive production traces. If it is a specialized decode path, it will.

Can I use it to replace my classifier?

Only for tasks that are already a closed set. Sentiment with four labels, ticket routing, "which tool next," "is this the same user or a new one." Not for "summarize the ticket" or "draft the reply." Those stay on Sol, Astra, or whatever writes your prose. The DevDay story is a split brain: Dots for long-running work, GPT-6.1 Sol for capable generation, Decisions API for the hop that should never have been a paragraph.

Honest limits (as of September 30, 2026)

  • Limited preview, selected customers. You cannot treat this as a platform default.
  • No published pricing. Any cost model you build today is a guess. Luna's public token prices are not a Decisions API rate card.
  • No public schema, SDK, or SLA. Latency quotes are stage and tweet, not a status page.
  • Still an LLM. Constraining Luna does not make it unable to pick a wrong label with high confidence. Calibration quality is unproven in public.
  • Answer-set design is on you. Garbage enums in, confident garbage out. "Other" as a bucket will absorb everything if you make it too attractive.
  • Visual inputs raise new failure modes. A screenshot of the wrong tab, a cropped receipt, a dark-mode contrast issue — the model will still return one of your answers.
  • Not a System One model. If your compliance story depended on "this component cannot generate text," Decisions API does not give you that invariant.

What this means for builders

If you are on the preview list, pick one production hop that is already an enum: support routing, tool selection, or "should we escalate." Log the current LLM-plus-schema baseline for a week (latency, cost, override rate). Swap only that hop. Keep generation on Sol or Astra. Do not rewrite the agent.

If you are not on the list, do not wait. Ship the same contract yourself: question, answers, context, selection, confidence, override. When Decisions API opens, the adapter is a weekend. That is also how you stay able to fail over to Jev or a local classifier without rewriting product logic.

If you were about to adopt Jev only because OpenAI had no equivalent, re-read the architecture row. You might still want Jev for the non-generative guarantee, for calibration research, or for a vendor that is not OpenAI. You might want Decisions API for vision + OpenAI ops. Those can coexist. They should not be the same Terraform module pretending they are one model.

Related on explainx.ai

  • OpenAI DevDay 2026: every announcement
  • GPT-6.1 Sol launch, pricing, and benchmarks
  • OpenAI Dots: always-on agents at DevDay
  • GPT-6 Sol and Luna launch pricing
  • How to integrate Jev for agent routing
  • Jev Ultrafast: browser-use with typed decisions
  • What is a System One model?
  • Official: OpenAI Developers announcement · DevDay · openai.com/live/

Figures and product claims are sourced to OpenAI Developers, OpenAI's DevDay recap wording, Tibo's latency comment, and public community reactions as of September 30, 2026. Decisions API is in limited preview; pricing and request schema were not published. Verify against OpenAI's live documentation before you ship.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 30, 2026

OpenAI Dots: Always-On GPT-6 Astra Agents Ship at DevDay 2026

On September 29, 2026, OpenAI shipped Dots — always-on GPT-6 Astra agents with their own cloud computer, browser, and 4,000+ app plugins. The first dot is included for eligible Pro and Business Premium users. Chat with a dot does not burn ChatGPT usage; Codex and Work tasks it starts do. The Live-from-the-Lab film is official; the on-stage demos did not all land.

Sep 29, 2026

OpenAI DevDay 2026: Every Announcement, Explained for Builders

OpenAI made 20+ announcements at DevDay on September 29, 2026. The headline is Dots, always-on ChatGPT agents with their own cloud computer, but the builder-relevant news is GPT-6.1 Sol pricing, Ultrafast, computer use in the Agents API, and plugin extensions. This is the explainx.ai recap, plus what it changes if you build on Claude.

Sep 27, 2026

Leaked ChatGPT Config Points at Always-On Consumer Agent "o"

On September 25–26, 2026, TestingCatalog reported unreleased ChatGPT client configuration referencing a lowercase-branded always-on consumer agent — display name "o," dedicated -o email suffixes, and UI copy consistent with proactive background assistance. OpenAI has not announced the product. DevDay is Tuesday, September 29, 2026 at Fort Mason. Here is what the leak actually shows, what is confirmed elsewhere, how it stacks against Meta Muse's shipped personal agent, and what builders running their own always-on stacks should prepare for.