On September 29 (US time), OpenAI announced the "Decisions API" and began offering it as a limited preview. The API lets developers define a question and several answer options, then chooses the appropriate one. It uses GPT-6 Luna and can be applied to tasks such as classifying content, routing requests, and selecting the next action for an AI agent.

Mechanisms that restrict generative AI output to set formats or options have existed before. What OpenAI is promoting this time is a dedicated API for handling, with low latency, the "decisions" that arise repeatedly inside apps and business systems.

If decisions become faster, it becomes easier to build AI into a software's branching logic. But choosing quickly is not the same as being able to trust those decisions with automated processing. For real-world use, what matters is not only speed but also how to verify the accuracy and reliability of the decisions.

AD

How does it differ from the existing Structured Outputs?

According to OpenAI Developers' announcement, developers set a question and answer options, and can build real-time decision-making into their apps.

For example, in routing customer inquiries, you could prepare options such as "shipping," "refund," "account," and "needs review," have the system read the text of an inquiry, and let it choose the right destination. This is an illustration of how the mechanism works, not an actual deployment example announced by OpenAI. Forwarding or replying after routing can be left to ordinary programs or another generative model.

That said, OpenAI already had a way to make a model answer from a fixed set of options.

Structured Outputs is a feature that makes a model's output follow a format defined with JSON Schema. You can specify required fields, or restrict values to a predefined list of options.

In other words, a task such as "return, in a fixed format, whether this inquiry is about a refund or shipping" can already be implemented with existing generation APIs.

So how does the Decisions API differ from existing features? Organizing the uses described in public materials gives the following:

Method What developers set Main role described in public materials
Structured Outputs Instructions for generated content and a JSON Schema Makes model output follow a specified format
Decisions API A question and answer options Classifies content and chooses a destination or next action
Moderation API Text or images to be evaluated Returns judgments and scores for categories defined by OpenAI

Divided by the roles in public materials, Structured Outputs is mainly a mechanism that constrains output format, the Decisions API is a mechanism that decides from a question and options set by the developer, and the Moderation API is a mechanism that makes judgments on categories predefined by OpenAI.

This table organizes the roles of each feature based on public materials available as of September 30, 2026. It is not a performance comparison under identical conditions.

Structured Outputs can also perform classification limited to set options. The Decisions API therefore does not add the ability to classify for the first time.

What is new is that it offers, as an independent API that can be called repeatedly with low latency, a process in which a question and answer options are given and a decision is made.

The Moderation API, for its part, has long been used to return judgment results for an input rather than generate long text. The Decisions API can be seen as extending this "return only a judgment" approach to the various business decisions that developers define themselves.

How to read the 150-millisecond figure

OpenAI's Tibo (@thsottiaux) explained that the Decisions API also supports image input and has been optimized to keep the time from input to decision result to under a few hundred milliseconds.

Because images of documents and screens can be used as decision material as well as text, uses are not limited to routing text. However, the post does not show how much performance varies depending on image size or content.

A more specific figure was reported by The New Stack, which also contacted OpenAI's public relations team.

Citing OpenAI's explanation, the outlet reports that the Decisions API returns results in about 150 milliseconds, whereas using GPT-6 Luna in the usual way takes about 1.6 seconds.

Putting the units on the same scale,

1,600 ms ÷ 150 ms ≈ 10.7

so a simple calculation suggests it is about 10 times faster.

However, this is a reported figure based on OpenAI's explanation, not a benchmark measured by a third party under the same conditions.

Some points about the comparison conditions also remain unclear. Input length, number of answer options, whether images were included, the region communicated from, and load under concurrent access have not been disclosed. It is also unknown which reasoning settings were used on the regular Luna side, or whether 150 ms and 1.6 seconds are averages or medians.

Therefore, it cannot be read as meaning that results come back in 150 ms for any input, nor can we assume the same response time when used from Japan.

Still, if the time needed to decide how processing branches gets shorter, subsequent searches and tool executions can start sooner.

For example, if a process that only reads an inquiry and chooses the responsible department no longer has to wait for long text generation, the time until a final answer reaches the user could potentially be shortened.

However, database access and the processing executed at the routing destination take separate time. The response speed of the Decisions API alone and the processing time of the whole app need to be evaluated separately.

AD

Jev, the "AI that doesn't generate text," got there first

On September 15, TypeSafe AI announced early access to "Jev," a model specialized for decision processing.

According to the company, Jev does not generate text; it returns typed decision results that programs can handle directly, along with probabilities indicating confidence.

The design philosophy of returning results that software can use directly for conditional branching, rather than generating long explanations, overlaps with the use the Decisions API is aiming for.

However, there is no information to conclude that the underlying technology is the same.

TypeSafe explains that Jev uses a new model architecture and parallel sampling, as well as "RLCD," a reinforcement learning method aimed at calibrating the probabilities returned at decision time.

OpenAI, on the other hand, has disclosed only that the Decisions API uses GPT-6 Luna. There is no basis for inferring that it adopts the same training method or parallel processing as Jev.

Jev's stated response time is 70 to 500 milliseconds, but TypeSafe notes that it ran many of its published evaluations from a laptop on the US West Coast.

In addition, the company itself explains that its comparison demo showing the effect of parallel sampling used short, information-dense inputs, a condition that favors Jev.

Meanwhile, its workflow evaluation uses the average of predictions by GPT-6 Astra and Fable 5.1 as the reference value. This differs in nature from an evaluation in which humans assign correct labels and accuracy against them is measured.

Numbers from such differing conditions cannot be directly compared with OpenAI's stated figures to rank the two companies on speed or decision accuracy.

Even so, the two companies' directions share something in common. Uses of generative AI are expanding beyond conversations that return free-form text to small decision processes called over and over inside software.

However, how far those decisions can be automated is not determined by speed alone.

In practice, what matters is not speed but choosing correctly

OpenAI itself states clearly that with Structured Outputs, the content itself can still be wrong even when the format is correct.

Answering in a specified format and choosing the correct option are separate problems.

Suppose, for example, you classify inquiries as either "shipping" or "refund." Even if the output format is always correct, if a message requesting a refund is judged to be "shipping," subsequent business processing goes in the wrong direction.

The New Stack explains that the Decisions API returns a confidence level along with the decision result.

However, the public materials we could confirm do not reveal the detailed response format of the Decisions API, nor any results verifying how closely that confidence matches actual accuracy.

This is where "probability calibration" becomes important.

If you collect 100 decisions on which the model showed 90% confidence and roughly 90 of them turn out to be correct, confidence and accuracy can be considered well matched.

This is an example to explain probability calibration, not a result actually measured for OpenAI or Jev.

Even if confidence is provided, it is not appropriate to simply set a rule such as "process automatically because it's 90%."

First, you need to measure accuracy using the Japanese-language inquiries you actually handle and the borderline cases that are hard to classify. Then you decide what level of confidence allows automated processing and what level should be sent for human review.

One approach to prepare for difficult cases is to include "needs review" among the answer options.

Averages alone are also not enough for speed. In addition to typical response times, checking how often processing is greatly delayed lets you evaluate the waiting time real users feel more accurately.

Regarding general availability, The New Stack reports that access will be expanded within the next few days, and that OpenAI's public relations team said it would release more details at that time.

Dedicated pricing, the upper limit on the number of answer options, and whether tuning with your own data is possible are not clear from the materials currently available.

Nor can we conclude from the explanation that it uses GPT-6 Luna that pricing will be the same as the regular Luna API.

If you compare once general availability begins, measure accuracy using the same inputs and the same answer options, and check how closely the confidence the model returns matches actual accuracy. Then compare response time and pricing.

Once these conditions are met, developers can separate decisions that tolerate some error from those that require human confirmation.

The value of the Decisions API is not simply making AI answers faster. It lies in being able to run small decisions, such as routing inquiries or choosing an AI agent's next action, repeatedly and with low latency inside software.

What determines its practicality is, in addition to that speed, how accurately it chooses and whether processing can be stopped safely when it gets something wrong.