OpenAI's Decisions API enters public beta

“OpenAI Decisions API” in white on a black background, surrounded by flowing strands of white, pale blue and amber particles.
In this article

OpenAI has opened its Decisions API to public beta: a dedicated endpoint that returns probabilities, choices and rubric scores instead of generated text, about 10x faster than the Responses API. Here is what it exposes and where it fits in agent and automation pipelines.

OpenAI's Decisions API is now in public beta. It evaluates text, images or both via a dedicated /v1/decisions endpoint and returns typed answers: a predicate probability from 0 to 1, a choice from a fixed set, or a score from the probability-weighted average of ordered levels. OpenAI documents it as about 10x faster than the Responses API, with gpt-6-luna the only model in the beta. It targets classification, routing and prioritization.

Classification, routing and triage are among the most common jobs teams hand to language models, and until now they usually meant treating a generation endpoint as a decision engine: prompt the model, parse the free text or JSON that comes back, and validate the shape yourself. OpenAI is now offering a narrower surface for exactly this kind of work. The company has put a Decisions API into public beta, documented in its developer platform guides, and the launch quickly drew attention on Hacker News, where it collected around 250 points and more than a hundred comments.

What the Decisions API exposes

According to OpenAI's documentation, the Decisions API evaluates text, images or both and returns typed answers about 10x faster than the Responses API. It runs on a dedicated POST /v1/decisions endpoint, and during the beta the only supported model is gpt-6-luna. OpenAI states that general availability is expected in the coming weeks, so the surface may still change before it stabilizes.

A request has three parts. The model field names the model that evaluates the request. The input carries the shared evidence for all questions, either as a text string or as user messages containing text and images. The questions array describes what to evaluate, including each question's type, its instructions and any allowed choices or score levels. The response contains an answers array in which each entry is identified by the unique name you gave the corresponding question, since the API echoes that name back.

Three question types, three answer shapes

The API offers three question types, each producing a different answer shape. The documentation is explicit about which one to reach for.

  • predicate: checks whether a condition is true, such as visible damage in a photo or the relevance of a passage, and returns a probability estimate from 0 to 1.

  • choice: selects one option from a fixed set you supply, such as a department or a content category, and returns one of those values.

  • score: rates an input against ordered levels, such as issue severity, and returns the probability-weighted average of the level indices.

The distinction between the last two matters in practice. Both choice and score return probabilities over discrete options, but choice is for categories without an order, while score is for ordered levels. Because a score is the probability-weighted average of numeric indices, the result can fall between levels, which suits gradual scales like severity better than a forced pick.

A worked example: checking a product photo

The guide's main example is a predicate question that inspects a product photo for visible damage. The input combines a text instruction with a base64-encoded image, and the question asks whether the product has a crack, tear or dent while explicitly telling the model to ignore shadows and damage to the packaging. The same pattern extends to text: the guide illustrates relevance checks on customer messages, where a clearly related request scores 1.00 and a question about returns after 30 days scores 0.90 against a policy answer.

JavaScript
import { readFile } from "node:fs/promises";
import OpenAI from "openai";

const client = new OpenAI();
const imageBase64 = (await readFile("product.png")).toString("base64");

const decision = await client.decisions.create({
  model: "gpt-6-luna",
  input: [
    {
      role: "user",
      content: [
        { type: "input_text", text: "Inspect the product in this photo." },
        { type: "input_image", image_url: `data:image/png;base64,${imageBase64}` },
      ],
    },
  ],
  questions: [
    {
      type: "predicate",
      name: "visible_damage",
      instructions:
        "Does the product have visible damage, such as a crack, tear, or dent? " +
        "Ignore shadows and damage to the packaging.",
    },
  ],
});

const answer = decision.answers[0];
if (answer.type === "refusal") {
  console.log(`Refused: ${answer.name}`);
} else if (answer.type === "predicate") {
  console.log(`Visible damage probability: ${answer.probability}`);
}

One detail worth noting for production code: the answers array can contain a refusal. The SDK examples check the answer type before reading a probability, which is a small but real difference from treating every model reply as usable output.

How it differs from completions and Structured Outputs

OpenAI draws the boundary clearly. Use Decisions when your application needs one of the three answer types above. Use Structured Outputs with the Responses API when you need to generate an object that follows your own JSON schema, such as extracted fields or a written explanation. Use function calling when you need the model to request a tool call with arguments. In other words, Decisions is not a replacement for generation; it is a specialized endpoint for questions whose answers are probabilities, categories or scores.

The practical consequence is architectural. A pipeline that today asks a chat model to classify a ticket and reply in JSON can move the classification step to a typed endpoint and reserve generation for the steps that actually need prose. That separation also makes decision points easier to test, since each question has a name, a type and a measurable answer rather than a blob of text to parse.

Where it fits in agent and automation architectures

The documented use cases are classification, request routing and prioritization of work. In an agent stack, these are exactly the frequent, latency-sensitive decision points that sit between more expensive generation steps: deciding which workflow handles an incoming message, grading the severity of an issue against a rubric, or gating a human review step on a probability threshold.

  • Support triage: score incoming tickets against severity levels and route the highest scores to humans first.

  • Content operations: classify documents or images into a fixed set of categories before storage or review.

  • Quality gates: use a predicate as a cheap check before committing to a longer, more expensive generation step.

  • Relevance filtering: measure whether a user question is actually answered by a knowledge base article, as in the documented examples.

Because the input is shared evidence for all questions, a single request can ask several things at once, for example whether an image shows damage, which category the issue belongs to and how severe it is, without three separate generation calls. The answers come back as a named array that maps directly onto branch logic in code.

SDK support and getting started

The feature is available through the official SDKs, with minimum versions listed in the guide: Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0 and Java 4.78.0. OpenAI also suggests trying the API in the Playground to experiment with questions and inputs before writing code. Authentication follows the standard pattern of an API key exported as an environment variable, which the SDKs read automatically.

Data handling notes

Anything you evaluate flows through the OpenAI platform, so the usual data controls apply. Data sent to the OpenAI API is not used to train or improve models unless you explicitly opt in, a policy in place since March 1, 2023. Abuse monitoring logs are generated by default and retained for up to 30 days, and customers who qualify can apply for Zero Data Retention or Modified Abuse Monitoring, subject to approval and additional requirements.

For teams evaluating the beta, the sensible reading is that Decisions is an infrastructure primitive rather than a product feature. It will matter most where decisions are high-volume, latency-sensitive and repeatable, and least where the answer needs explanation or nuance. The single-model limitation and the beta label both argue for piloting on internal workflows before moving customer-facing paths onto it.

How to evaluate the beta in practice

  1. Pick one internal workflow with a clear decision point, such as ticket routing or image screening, and express its questions as predicates, choices or scores.

  2. Run historical cases through the Playground and the API, and compare the probabilities and scores against your current labels or human decisions.

  3. Set explicit thresholds for acting on probabilities, and log refusals separately from answers so they can be handled as their own case.

  4. Keep a fallback path through the Responses API with Structured Outputs until the API reaches general availability and the model lineup expands.

The Decisions API compresses a pattern that many teams have built by hand: constrained questions, typed answers and fast evaluation. Whether it earns a permanent place in agent architectures will depend on how the beta behaves under real traffic, but the surface it exposes is easy to understand and straightforward to benchmark against what you already run.

Key takeaways

  • The Decisions API is in public beta with general availability expected in the coming weeks, documented in OpenAI's developer guides.
  • It evaluates text, images or both and returns typed answers about 10x faster than the Responses API, according to OpenAI's documentation.
  • Three question types are available: predicate for a 0 to 1 probability, choice for one of a fixed set of options, and score for ordered levels.
  • Requests go to a dedicated POST /v1/decisions endpoint, and gpt-6-luna is currently the only supported model.
  • Each question carries a unique name that the API echoes in the answers array, and answers can also come back as refusals.
  • OpenAI positions Decisions for classification, routing and prioritization, while Structured Outputs and function calling cover JSON generation and tool calls.
  • Newer SDK versions are required: Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0 and Java 4.78.0.
  • API data is not used for training by default, and abuse monitoring logs are retained up to 30 days unless Zero Data Retention or Modified Abuse Monitoring controls apply.

Frequently asked questions

What is the OpenAI Decisions API?

It is a dedicated API, currently in public beta, that evaluates text, images or both and returns typed answers instead of generated prose. It supports three question types: predicate, which returns a probability from 0 to 1; choice, which returns one of a supplied set of values; and score, which rates an input against ordered levels.

How fast is the Decisions API compared to the Responses API?

OpenAI's documentation states that the Decisions API returns typed answers about 10x faster than the Responses API. The company positions it for high-volume decision points such as classification, request routing and work prioritization.

Which models can I use with the Decisions API?

During the public beta, gpt-6-luna is the only model available. Requests are sent to the dedicated POST /v1/decisions endpoint.

When should I use Decisions instead of Structured Outputs?

Use Decisions when your application needs a probability, a choice from a fixed set or a rubric score. Use Structured Outputs with the Responses API when you need to generate an object that follows your own JSON schema, such as extracted fields or a written explanation, and function calling when the model needs to request a tool call with arguments.

Which SDK versions support the Decisions API?

The documentation lists minimum versions of Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0 and Java 4.78.0. You can also try questions and inputs in the Playground before writing code.

What happens to the data I send to the Decisions API?

Data sent to the OpenAI API is not used to train or improve models unless you explicitly opt in. Abuse monitoring logs are generated by default and retained for up to 30 days; approved customers can apply for Zero Data Retention or Modified Abuse Monitoring controls.

Can the Decisions API refuse to answer?

Yes. The answers array includes a refusal type, and the SDK examples show code checking for a refusal before reading a predicate probability, so refusals should be handled as their own case in production logic.

Sources

More on this topic

Digital Strategy
Digital Strategy
Software Architecture
Software Architecture
Development
Development
User Experience
User Experience
Mobile Apps
Mobile Apps
Artificial Intelligence
Artificial Intelligence
Cybersecurity
Cybersecurity
Automation
Automation
Cloud Infrastructure
Cloud Infrastructure
DevOps
DevOps
Digital Strategy
Digital Strategy
Software Architecture
Software Architecture
Development
Development
User Experience
User Experience
Mobile Apps
Mobile Apps