When should you use OpenAI Decisions instead of Responses?
AI Architect
Key takeaways
- The Decisions API accepts only `predicate`, `choice`, and `score` questions and returns typed answers in an `answers` array.
- OpenAI's Decisions guide says image inputs must be inline base64 data URLs, so hosted image URLs and `file_id` inputs fail on this endpoint.
- OpenAI's pricing for Decisions with `gpt-6-luna` charges input tokens only, while the same model on Responses also has cached-input, cache-write, and output-token rates.
- OpenAI's score questions return a probability-weighted average over ordered levels, so a severity result can land between labels rather than on one label.
- Structured Outputs and function calling stay the better fit whenever an application needs extracted fields, written rationale, or a real tool loop.
What did OpenAI actually ship in the Decisions API?
OpenAI’s 6 October 2026 changelog says the new Decisions API turns text and
images into typed answers on a dedicated POST /v1/decisions endpoint and runs
about 10 times faster than Responses for that job class. The guide then defines
the boundary more tightly than most summaries did. A request has model,
input, and questions; the response comes back as an answers array keyed by
question name. The only question types are predicate, choice, and
score, and the only model in beta is gpt-6-luna.
That matters because this is a scorer for bounded questions. OpenAI even
routes readers away from it when the output needs to follow a custom JSON
schema or request a tool call. One practical constraint is easy to miss in a
launch skim: image inputs must be inline base64 data URLs. The guide says
hosted HTTP or HTTPS image URLs and file_id image inputs are unsupported, so
an existing image pipeline may need a small rewrite before a team sees any of
the speed or cost benefits.
Where the Decisions surface stops
Show as text
| # | Layer | Note |
|---|---|---|
| 1 | Shared evidence in input | Text or mixed text plus inline images. |
| 2 | Typed question list | Predicate, choice, or score only. |
| 3 | answers array with probabilities | Endpoint returns typed results. |
| · | compare with Responses below (breakpoint) | |
| 4 | Prompt or message input | Same family, broader output surface. |
| 5 | JSON schema or tool definitions | Extraction and tool use live here. |
| 6 | Generated text or tool loop | Output can be prose, JSON, or calls. |
How is OpenAI Decisions different from Structured Outputs?
The cleanest split is answer shape. Decisions returns one of three built-in types: a truth estimate, a choice from your own fixed set, or a score across ordered levels. Structured Outputs on Responses does something broader. It asks the model to generate an object that conforms to your JSON schema, which is why it suits extraction, normalization, and any result that needs named fields the API does not predefine. That is a wider contract and a heavier one.
The score path is the detail worth carrying into design reviews. OpenAI defines
score as the probability-weighted average of ordered level indices, so a result
can land between labels. A severity score of 1.1 means the model spread
probability across adjacent levels. Teams that need hard buckets should still keep the
distribution or threshold band, much like the discipline described in
question design for classification and in
type-safe AI system one. My view is that Decisions
is the better first hop when the app only needs a typed gate. The moment you
need extracted fields or a written explanation, Responses has already won.
Why can Decisions be cheaper and faster than Responses?
OpenAI’s own pricing explanation gives the short answer: Decisions with
gpt-6-luna charges input tokens only. The guide says there are no cache-read,
cache-write, or output-token charges on /v1/decisions. By contrast, OpenAI’s
pricing page lists multiple billable surfaces for gpt-6-luna on Responses,
including input, cached input, cache writes, and output tokens. Even before you
benchmark latency, the contract tells you why a typed scorer can be quicker and
operationally simpler than a generation endpoint.
The cost story still needs discipline. Responses can earn back some of that extra surface with prompt caching, which matters in high-reuse traffic and is a big reason to keep reading work such as prompt caching economics. Decisions also gives up expressive power to get its speed. There is no written rationale for a human reviewer, no schema-shaped extraction result, and no tool loop. So the real question is not whether Decisions is cheaper in a vacuum. The question is whether the application can stop at a typed answer without paying a second request to another surface.
What GPT-6 Luna bills on the two surfaces
Show data
| Item | Value (USD per 1M tokens) | Note |
|---|---|---|
| Decisions input | 0.1 | Decisions guide says input only. |
| Responses input | 0.1 | Standard short-context input rate. |
| Responses cached input | 0.01 | Applies when a cache hit exists. |
| Responses output | 0.5 | Generated output has its own rate. |
| Responses cache write | 0.125 | Prefix writes are billed separately. |
Which workloads fit OpenAI Decisions best today?
The strongest fits are first-hop routing, triage, and review gating. OpenAI’s own examples are visible damage checks, department routing, and severity scoring. Those are all jobs where the application wants a bounded answer and already knows what to do next. A moderation queue can route anything above a predicate threshold to manual review. A support desk can choose a department. A bug intake service can score severity, then send only borderline tickets to a full Responses flow for explanation or follow-up.
This is also where the data-controls note matters. OpenAI says Decisions is ZDR eligible and supports HIPAA use for eligible customers, with regional processing in the United States and Europe. That makes it easier to justify for production triage pipelines than a brand-new endpoint would normally be. My opinion is that Decisions works best as a front door for a larger workflow. Full case handling belongs elsewhere. If a team keeps the boundary narrow and preserves an escalation path, it can cut cost and latency without pretending a typed score is the same thing as a reviewed decision.
When should you stay on OpenAI Responses instead?
Stay on Responses when the work product must say more than a bounded answer. OpenAI’s Decisions guide itself points readers to Structured Outputs when they need extracted fields or written explanation, and to function calling when the model must request a real action. In those workloads, the larger control surface matches the work itself. If the app needs a refund tool call, a schema-shaped claim form, or a rationale a human can audit, Responses is still the right surface even when it costs more.
There is another practical reason to stay put: migration friction. Teams with
existing image pipelines on hosted URLs or file_id inputs cannot drop those
requests straight onto Decisions because the guide requires inline base64 image
data. That conversion may be fine for a hot path and annoying for a batch path.
My view is simple. Use Decisions for high-volume gates where the next action is
already determined by the answer type. Use Responses everywhere the answer still
needs to explain itself, extract a richer object, or touch the outside world.
Do this
Choose between Decisions and Responses
Start from the shape of the answer your application actually needs, then test the failure mode before you optimize cost.
Check whether the output can be only predicate, choice, or score
If the application needs a paragraph, extracted object, or tool request, stop and use Responses instead of trying to coerce Decisions into a wider job.
Define choices or score levels with a fallback path
Add an `other` category or a review queue so the model has a safe place to put ambiguous inputs instead of forcing a false certainty.
Set thresholds from labelled examples
Decisions returns probabilities and confidence. Pick thresholds by replaying real samples and pricing the cost of false accepts against false reviews.
Escalate uncertain cases to Responses
Use Decisions as the first gate, then send only low-confidence or high-consequence cases to Structured Outputs or a tool loop where richer reasoning is worth the extra surface.
Frequently asked questions
- Can OpenAI Decisions replace Structured Outputs for extraction?
- No. The Decisions guide limits the surface to predicate, choice, and score questions. OpenAI directs applications that need extracted fields or written explanations to Structured Outputs on the Responses API.
- Can OpenAI Decisions call tools?
- No. The Decisions guide points tool-using applications to function calling. Decisions returns typed answers only and does not run a tool loop.
- Why can a Decisions score return 1.1 instead of one severity label?
- OpenAI defines score as the probability-weighted average of ordered level indices. That means the result can sit between levels and should be thresholded by your application.
- Can OpenAI Decisions read an HTTPS image URL?
- No. The guide says images must be sent as inline base64 data URLs. Hosted HTTP or HTTPS URLs and `file_id` image inputs are not supported on this endpoint.
- Which model supports OpenAI Decisions today?
- The Decisions guide says `gpt-6-luna` is the only supported model during public beta.
Sources
Revision history
- : First published against OpenAI's Decisions API beta release and documentation.
- : First published.