DeepThinking AI

When should you use OpenAI Decisions instead of Responses?

AI Architect

Key takeaways

  • The Decisions API accepts only `predicate`, `choice`, and `score` questions and returns typed answers in an `answers` array.
  • OpenAI's Decisions guide says image inputs must be inline base64 data URLs, so hosted image URLs and `file_id` inputs fail on this endpoint.
  • OpenAI's pricing for Decisions with `gpt-6-luna` charges input tokens only, while the same model on Responses also has cached-input, cache-write, and output-token rates.
  • OpenAI's score questions return a probability-weighted average over ordered levels, so a severity result can land between labels rather than on one label.
  • Structured Outputs and function calling stay the better fit whenever an application needs extracted fields, written rationale, or a real tool loop.

What did OpenAI actually ship in the Decisions API?

OpenAI’s 6 October 2026 changelog says the new Decisions API turns text and images into typed answers on a dedicated POST /v1/decisions endpoint and runs about 10 times faster than Responses for that job class. The guide then defines the boundary more tightly than most summaries did. A request has model, input, and questions; the response comes back as an answers array keyed by question name. The only question types are predicate, choice, and score, and the only model in beta is gpt-6-luna.

That matters because this is a scorer for bounded questions. OpenAI even routes readers away from it when the output needs to follow a custom JSON schema or request a tool call. One practical constraint is easy to miss in a launch skim: image inputs must be inline base64 data URLs. The guide says hosted HTTP or HTTPS image URLs and file_id image inputs are unsupported, so an existing image pipeline may need a small rewrite before a team sees any of the speed or cost benefits.

Where the Decisions surface stops

Where the Decisions surface stopsDiagram: 7 ordered layers. Shared evidence in input, then Typed question list, then answers array with probabilities, then compare with Responses below (breakpoint), then Prompt or message input, then JSON schema or tool definitions, then Generated text or tool loop.1Shared evidence in inputText or mixed text plus inline images.2Typed question listPredicate, choice, or score only.3answers array with probabilitiesEndpoint returns typed results.compare with Responses below4Prompt or message inputSame family, broader output surface.5JSON schema or tool definitionsExtraction and tool use live here.6Generated text or tool loopOutput can be prose, JSON, or calls.
Show as text
Where the Decisions surface stops. Diagram: 7 ordered layers. Shared evidence in input, then Typed question list, then answers array with probabilities, then compare with Responses below (breakpoint), then Prompt or message input, then JSON schema or tool definitions, then Generated text or tool loop.
#LayerNote
1Shared evidence in inputText or mixed text plus inline images.
2Typed question listPredicate, choice, or score only.
3answers array with probabilitiesEndpoint returns typed results.
·compare with Responses below (breakpoint)
4Prompt or message inputSame family, broader output surface.
5JSON schema or tool definitionsExtraction and tool use live here.
6Generated text or tool loopOutput can be prose, JSON, or calls.
Decisions is a typed scorer sitting beside the broader Responses surface. Its boundary is the answer type. The model family is shared.

How is OpenAI Decisions different from Structured Outputs?

The cleanest split is answer shape. Decisions returns one of three built-in types: a truth estimate, a choice from your own fixed set, or a score across ordered levels. Structured Outputs on Responses does something broader. It asks the model to generate an object that conforms to your JSON schema, which is why it suits extraction, normalization, and any result that needs named fields the API does not predefine. That is a wider contract and a heavier one.

The score path is the detail worth carrying into design reviews. OpenAI defines score as the probability-weighted average of ordered level indices, so a result can land between labels. A severity score of 1.1 means the model spread probability across adjacent levels. Teams that need hard buckets should still keep the distribution or threshold band, much like the discipline described in question design for classification and in type-safe AI system one. My view is that Decisions is the better first hop when the app only needs a typed gate. The moment you need extracted fields or a written explanation, Responses has already won.

Why can Decisions be cheaper and faster than Responses?

OpenAI’s own pricing explanation gives the short answer: Decisions with gpt-6-luna charges input tokens only. The guide says there are no cache-read, cache-write, or output-token charges on /v1/decisions. By contrast, OpenAI’s pricing page lists multiple billable surfaces for gpt-6-luna on Responses, including input, cached input, cache writes, and output tokens. Even before you benchmark latency, the contract tells you why a typed scorer can be quicker and operationally simpler than a generation endpoint.

The cost story still needs discipline. Responses can earn back some of that extra surface with prompt caching, which matters in high-reuse traffic and is a big reason to keep reading work such as prompt caching economics. Decisions also gives up expressive power to get its speed. There is no written rationale for a human reviewer, no schema-shaped extraction result, and no tool loop. So the real question is not whether Decisions is cheaper in a vacuum. The question is whether the application can stop at a typed answer without paying a second request to another surface.

What GPT-6 Luna bills on the two surfaces

What GPT-6 Luna bills on the two surfacesBar chart. Decisions input: 0.1 USD per 1M tokens. Responses input: 0.1 USD per 1M tokens. Responses cached input: 0.01 USD per 1M tokens. Responses output: 0.5 USD per 1M tokens. Responses cache write: 0.125 USD per 1M tokens.Decisions input0.1 USD per 1M tokensDecisions guide says input only.Responses input0.1 USD per 1M tokensStandard short-context input rate.Responses cached input0.01 USD per 1M tokensApplies when a cache hit exists.Responses output0.5 USD per 1M tokensGenerated output has its own rate.Responses cache write0.125 USD per 1M tokensPrefix writes are billed separately.
Show data
What GPT-6 Luna bills on the two surfaces. Bar chart. Decisions input: 0.1 USD per 1M tokens. Responses input: 0.1 USD per 1M tokens. Responses cached input: 0.01 USD per 1M tokens. Responses output: 0.5 USD per 1M tokens. Responses cache write: 0.125 USD per 1M tokens.
ItemValue (USD per 1M tokens)Note
Decisions input0.1Decisions guide says input only.
Responses input0.1Standard short-context input rate.
Responses cached input0.01Applies when a cache hit exists.
Responses output0.5Generated output has its own rate.
Responses cache write0.125Prefix writes are billed separately.
This is a view of billable dimensions rather than a full request total. OpenAI lists input-only billing for Decisions on GPT-6 Luna, while Responses can bill input, cached input, cache writes, and output because it supports generated text and prompt-cache state.

Which workloads fit OpenAI Decisions best today?

The strongest fits are first-hop routing, triage, and review gating. OpenAI’s own examples are visible damage checks, department routing, and severity scoring. Those are all jobs where the application wants a bounded answer and already knows what to do next. A moderation queue can route anything above a predicate threshold to manual review. A support desk can choose a department. A bug intake service can score severity, then send only borderline tickets to a full Responses flow for explanation or follow-up.

This is also where the data-controls note matters. OpenAI says Decisions is ZDR eligible and supports HIPAA use for eligible customers, with regional processing in the United States and Europe. That makes it easier to justify for production triage pipelines than a brand-new endpoint would normally be. My opinion is that Decisions works best as a front door for a larger workflow. Full case handling belongs elsewhere. If a team keeps the boundary narrow and preserves an escalation path, it can cut cost and latency without pretending a typed score is the same thing as a reviewed decision.

When should you stay on OpenAI Responses instead?

Stay on Responses when the work product must say more than a bounded answer. OpenAI’s Decisions guide itself points readers to Structured Outputs when they need extracted fields or written explanation, and to function calling when the model must request a real action. In those workloads, the larger control surface matches the work itself. If the app needs a refund tool call, a schema-shaped claim form, or a rationale a human can audit, Responses is still the right surface even when it costs more.

There is another practical reason to stay put: migration friction. Teams with existing image pipelines on hosted URLs or file_id inputs cannot drop those requests straight onto Decisions because the guide requires inline base64 image data. That conversion may be fine for a hot path and annoying for a batch path. My view is simple. Use Decisions for high-volume gates where the next action is already determined by the answer type. Use Responses everywhere the answer still needs to explain itself, extract a richer object, or touch the outside world.

Do this

Choose between Decisions and Responses

Start from the shape of the answer your application actually needs, then test the failure mode before you optimize cost.

  1. Check whether the output can be only predicate, choice, or score

    If the application needs a paragraph, extracted object, or tool request, stop and use Responses instead of trying to coerce Decisions into a wider job.

  2. Define choices or score levels with a fallback path

    Add an `other` category or a review queue so the model has a safe place to put ambiguous inputs instead of forcing a false certainty.

  3. Set thresholds from labelled examples

    Decisions returns probabilities and confidence. Pick thresholds by replaying real samples and pricing the cost of false accepts against false reviews.

  4. Escalate uncertain cases to Responses

    Use Decisions as the first gate, then send only low-confidence or high-consequence cases to Structured Outputs or a tool loop where richer reasoning is worth the extra surface.

Frequently asked questions

Can OpenAI Decisions replace Structured Outputs for extraction?
No. The Decisions guide limits the surface to predicate, choice, and score questions. OpenAI directs applications that need extracted fields or written explanations to Structured Outputs on the Responses API.
Can OpenAI Decisions call tools?
No. The Decisions guide points tool-using applications to function calling. Decisions returns typed answers only and does not run a tool loop.
Why can a Decisions score return 1.1 instead of one severity label?
OpenAI defines score as the probability-weighted average of ordered level indices. That means the result can sit between levels and should be thresholded by your application.
Can OpenAI Decisions read an HTTPS image URL?
No. The guide says images must be sent as inline base64 data URLs. Hosted HTTP or HTTPS URLs and `file_id` image inputs are not supported on this endpoint.
Which model supports OpenAI Decisions today?
The Decisions guide says `gpt-6-luna` is the only supported model during public beta.

Sources

  1. API changelogOpenAI · 2026-10-06
  2. DecisionsOpenAI
  3. Structured model outputsOpenAI
  4. Function callingOpenAI
  5. PricingOpenAI
  6. Data controls in the OpenAI platformOpenAI

Revision history

  • : First published against OpenAI's Decisions API beta release and documentation.
  • : First published.

openaidecisions-apiresponses-apistructured-outputsrouting