OpenAI’s Decisions API Picks Answers Instead of Writing Them

OpenAI introduced the Decisions API at DevDay on September 29, 2026 — an endpoint that does not write text at all. Developers define a question and a finite set of allowed answers, pass in text or an image as context, and get one of those answers back. It runs on a specialised version of GPT‑6 Luna, and OpenAI says it returns a decision in about 150 milliseconds, against roughly 1.6 seconds for a standard Luna call. It launched in limited preview, with a broad release “planned in the coming days.” It arrives two weeks after TypeSafe AI launched Jev, a model built around the same idea.
Intermediate
What OpenAI Announced
OpenAI’s DevDay recap describes the product in one paragraph: the Decisions API “enables real-time decision-making by focusing Luna’s intelligence on a specific set of user-defined questions with finite pre-defined answers.” Developers supply context as text or images and get back answers they can use to “classify content, route requests, or choose an agent’s next action.”
The headline number is latency. On OpenAI’s slide, a Decisions API call completes in 150 ms and a GPT‑6 Luna API call in 1.6 s — the basis for the “10x faster” claim. These are OpenAI’s figures. No independent measurement has been published yet, and OpenAI has not said what task or input size the comparison used.
The Decisions API was one of more than 20 DevDay announcements. The others included GPT‑6.1 Sol, the Ultrafast speed tier, computer use in the Agents API, and Dots, OpenAI’s always-on agents.
Why a Closed Answer Set Matters
A large share of production LLM calls are not really generation tasks. Examples include checking whether a support ticket is a billing question, deciding whether an image breaks a content policy, or choosing which tool an agent calls next. Developers usually handle these by prompting a chat model, asking for JSON, and parsing the reply. Structured Outputs made that reply reliably parseable, but the model still generates tokens to produce it, and the calling code still has to handle an answer that is valid JSON but not one of the expected labels.
The Decisions API removes that step. The answer set is part of the request rather than something enforced after generation, so there is nothing free-form to parse or validate. As FourWeekMBA puts it, the API limits the output space itself instead of filtering outputs after the fact. The latency gain follows from the same design: returning one answer from a known list takes far less decoding than writing a sentence that contains it.
What Has Not Been Published
Much of what developers need to evaluate the product is still missing. As of September 30:
- No documentation. The Decisions API does not appear in OpenAI’s API changelog, which lists other DevDay releases such as GPT‑6.1 Sol and computer use in the Agents API.
- No pricing. Standard GPT‑6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens, but OpenAI has not said whether Decisions calls are billed the same way. OrcaRouter cautions that any per-decision cost quoted now is an inference.
- No stated limits or accuracy figures. OpenAI has not said how many candidate answers a request can carry, whether the response includes a probability or confidence value for each option, or how accurate the model is on standard classification benchmarks.
What This Means
This is the second “decision model” launch in a month. TypeSafe AI’s Jev came out of stealth on September 15. It uses the same interface — typed questions and developer-defined answers — and returns a probability for every option. TypeSafe says Jev’s end-to-end latency is 70–500 ms and prices it at $0.042 per million input tokens, with output free. OpenAI’s version accepts images as well as text and comes with the rest of OpenAI’s platform, but it has not yet published the numbers Jev’s launch included.
Taken together, the two launches suggest where agent design is heading. An agent loop has a planner that reasons in free text and many small routing steps. The planner needs a frontier model; the routing steps need a fast, cheap answer from a closed list. Using a full chat model for every small step has been expensive and slow. Splitting those steps onto a dedicated endpoint could cut both cost and latency, if the accuracy holds up on real data.
For teams already using prompt-and-parse classifiers, the practical advice is to wait for the documentation. The two things to check are whether the API returns per-option probabilities, which is what makes thresholds, abstention, and escalation to a larger model possible, and whether the Decisions pricing keeps Luna’s rates.
Related Coverage
- TypeSafe AI Launches Jev, a Model That Returns Decisions, Not Text — the system-one model the Decisions API is being compared to
- OpenAI Releases GPT-6 Sol and Luna at Half the Price — the launch of the model the Decisions API is built on
- OpenAI Launches Dots, Always-On Agents Powered by GPT-6 Astra — another DevDay 2026 announcement
- OpenAI Introduces Structured Outputs in the API — OpenAI’s earlier approach to constrained output
This post was drafted with AI assistance and reviewed by RITS staff.
Sources
- OpenAI — DevDay 2026 Recap (September 29, 2026)
- The Decoder — OpenAI expands Codex and its API at DevDay with security scans, a Decisions API, and Ultrafast
- FourWeekMBA — OpenAI Decisions API Limits the Output Space Itself
- OrcaRouter — OpenAI’s Decisions API: GPT-6 Luna Picks One Answer
- OpenAI API Changelog



沪公网安备31011502017015号