Should you use OpenAI background mode or the Batch API?
AI Architect
Key takeaways
- OpenAI background mode keeps work inside one Responses object that your application can poll, stream, or cancel.
- OpenAI Batch can also call `/v1/responses`, so choosing Batch does not mean giving up the Responses API surface.
- OpenAI Batch gives 50% lower cost and a separate pool of significantly higher rate limits, but only with a `24h` completion window.
- OpenAI background mode may delete response data after roughly 10 minutes unless retention is explicitly requested where the policy allows it.
- OpenAI Batch writes results to output and error files keyed by `custom_id`, which changes how retries and result joins work.
The most useful line in OpenAI’s current async docs is buried in the Batch
guide rather than the background guide: Batch accepts /v1/responses as an
input endpoint. That turns the decision into a control-plane choice over the
same underlying model surface. For teams building long-running AI jobs, that is
the distinction worth optimising.
My view is straightforward. Start with background mode when a single job still belongs to your application flow, especially on agent backends where progress, cancellation, and per-job debugging matter. Move to Batch when the work has become a backlog problem and the 50 percent discount plus higher rate-limit pool matter more than immediacy. That is a cleaner split than most launch summaries give it.
What does OpenAI background mode actually preserve?
Background mode preserves the normal Responses object as the unit your
application talks to. You create a request with background: true, receive a
response id back, then keep polling while status is queued or
in_progress. The same guide documents response-level cancellation, and it also
documents a hybrid path where you set both background and stream to true so
the job continues even if the client drops the connection. That is a real
control feature rather than a cosmetic wrapper around long latency.
The engineering implication is that your existing job model can often stay intact. One request becomes one tracked object with one cancellation path and one place to retrieve output. If you are already working with hosted OpenAI job surfaces such as the Agents API runtime boundary, background mode fits naturally because it behaves like a live job inside your application. It feels very different from a later file import.
Where the two async paths diverge
Show as text
| # | Layer | Note |
|---|---|---|
| 1 | Responses request with background true | One job stays a normal response object. |
| 2 | Poll, stream, or cancel by response id | Application keeps per-job control. |
| 3 | response object carries final output | Read it back through Responses. |
| · | compare with Batch below (breakpoint) | |
| 4 | JSONL file of many requests | Each line needs a custom_id. |
| 5 | 24 hour batch worker window | Higher throughput, lower urgency. |
| 6 | output and error files | Results arrive through files and `custom_id`. |
What does the OpenAI Batch API add that background mode does not?
Batch adds three guarantees the background guide does not make: 50 percent lower
cost than synchronous APIs, a separate pool of significantly higher rate limits,
and a 24h completion window. It also changes the input contract. Work starts
as a .jsonl file, each request line carries its own custom_id, and OpenAI
processes the set as one batch object rather than as individually retrievable
response ids. Even when the underlying request body targets /v1/responses, the
job-control layer is different.
That difference is why Batch suits offline operations better than interactive ones. Evaluations, bulk classification, repository embeddings, and queued content transformations usually want price and throughput more than they want a stream that can resume mid-flight. The file boundary also makes ownership cleaner for operations teams that already think in staged artifacts, similar to the queue-first runtime pattern in cloud agent deployment. What you lose is the per-request feel of a live application call.
Relative price index for the two async paths
Show data
| Item | Value (index) | Note |
|---|---|---|
| Responses request, sync or background | 100 | baseline pricing path |
| Batch API request | 50 | guide states 50 percent lower cost |
How do OpenAI background mode and Batch differ on retention?
This is the detail most teams miss, and it is the strongest reason to read the
primary docs rather than secondary tutorials. OpenAI’s background guide and
“Your data” guide both say background responses may be deleted after roughly 10
minutes when store is omitted or false, with a narrower rule for projects
using Modified Abuse Monitoring: long-term retention after the polling period
happens only when store=true is explicitly provided where policy permits it.
So background mode is operationally retrievable, but it is not a durable
results store by default.
Batch behaves differently. The batch object points to output_file_id and, if
needed, error_file_id, and the guide says the output file is automatically
deleted 30 days after the batch completes. That is still temporary, but it is a
very different operational envelope from roughly 10 minutes. If a downstream
system joins results slowly, Batch gives you much more slack. If your design
needs permanent history, neither path replaces writing the result into your own
database after retrieval.
When should OpenAI background mode be your default?
Background mode should be your default when the important noun in the system is “job”, singular. A user submitted one research task. One internal service kicked off one agent run. One operator may need to cancel one response when a budget cap trips. In those situations, the response id is the right operational key, and the ability to poll, cancel, or resume a stream is more valuable than batch discounts. The control surface matches the mental model.
This is especially true for long-running agent work. An agent loop often wants the same application to watch progress, tie logs to one request, and stop work when a human intervenes. Batch can carry the same model request body, but the cost of unpacking results from files usually exceeds the savings until volume is high. My opinion is that teams reach for Batch too early because the 50 percent number is loud. Control-plane friction is quieter, but it is what operators pay for every day.
When should the OpenAI Batch API win?
Batch should win when the important noun is “backlog”. You have thousands of
requests, none of them need immediate delivery, and the surrounding system is
already comfortable with file upload, later collection, and result joins through
custom_id. That is the shape OpenAI describes directly in the guide: running
evals, classifying large datasets, and embedding content repositories. In those
cases, lower cost and more throughput are first-order advantages rather than
nice extras.
Batch is also the more honest fit when expiry is acceptable. The guide states that batches can expire, with unfinished requests recorded in the error file, and that any completed requests still appear in the output file. That partial delivery model is normal for offline processing and awkward for user-facing job control. If your consumer can reconcile a mixed output file and resubmit failed or expired lines, Batch is doing exactly the job it was designed for. If not, stay with background mode and keep the application loop simpler.
Do this
Choose between OpenAI background mode and the Batch API
Start with the application boundary, then move to urgency, price, and retention. That order removes the most expensive category errors first.
Ask whether one job must stay inside your live application flow
If another service, a user session, or an operator still needs one job id they can poll or cancel, start with background mode. It keeps the unit of work as a normal Responses object.
Check whether the work can tolerate a file-based 24 hour lane
Batch only offers a `24h` completion window and starts from a `.jsonl` upload. If that sounds normal for the job, such as nightly evals or mass classification, Batch is a candidate.
Compare price pressure and throughput pressure together
Batch wins when the backlog is large because OpenAI documents both a 50 percent discount and a separate pool of significantly higher rate limits. Background mode keeps ordinary request economics.
Decide how your system will read results back
Background responses are retrieved by response id and can also be streamed. Batch writes output and error files keyed by `custom_id`, so downstream joins and retries happen at the file layer.
Review retention and deletion rules before launch
Background mode may delete response data after roughly 10 minutes unless `store=true` is explicitly allowed and set. Batch output files persist much longer, but they are still temporary operational artifacts rather than a permanent database.
Pilot one representative workload before standardising
Run the same workload shape once through background mode and once through Batch. Measure operational friction alongside tokens, because the cheaper path can still be the wrong one if it breaks the surrounding system.
Frequently asked questions
- Can the Batch API call the Responses API?
- Yes. OpenAI's Batch guide lists `/v1/responses` as a supported endpoint. That means the background mode versus Batch decision is about delivery mechanics, rate limits, retention, and price rather than about a separate model feature set.
- Does background mode cost less than a normal Responses request?
- The background guide does not advertise a discount. The Batch guide does advertise 50% lower cost than synchronous APIs, so cost-sensitive offline work belongs there unless you need the single-response control surface.
- Can I cancel both kinds of work?
- Yes, but the cancellation unit is different. Background mode cancels one response object. Batch cancels the whole batch job, then waits for in-flight requests to finish for up to 10 minutes before the batch becomes cancelled.
- Which path is better for agent backends?
- Background mode is usually better when a user or another service still cares about one job's status, partial progress, or cancellation. Batch is better when the work is a large offline set such as evaluations, repository embeddings, or bulk classification.
- What is the easiest mistake to make when switching to Batch?
- Treating it like a cheaper background response. Batch changes the input shape to `.jsonl`, requires `custom_id` per request, and returns results through output files rather than a normal Responses retrieve call.