---
title: Should you use OpenAI background mode or the Batch API?
url: https://deepthinkingai.org/openai-background-mode-vs-batch-api/
published: 2026-10-01
author: Shekhar Singh
topic: AI Engineering
tags: openai, background-mode, batch-api, responses-api, async-jobs
site: DeepThinking AI
---

# Should you use OpenAI background mode or the Batch API?

**Summary:** Because OpenAI Batch can target `/v1/responses`, the real choice is the job control plane. Use `background: true` when one long-running job must stay pollable, streamable or cancellable inside your app. Use Batch when 50% lower cost, higher rate limits and a 24 hour file workflow matter more than immediacy.

## Key takeaways
- OpenAI background mode keeps work inside one Responses object that your application can poll, stream, or cancel.
- OpenAI Batch can also call `/v1/responses`, so choosing Batch does not mean giving up the Responses API surface.
- OpenAI Batch gives 50% lower cost and a separate pool of significantly higher rate limits, but only with a `24h` completion window.
- OpenAI background mode may delete response data after roughly 10 minutes unless retention is explicitly requested where the policy allows it.
- OpenAI Batch writes results to output and error files keyed by `custom_id`, which changes how retries and result joins work.

The most useful line in OpenAI's current async docs is buried in the Batch
guide rather than the background guide: Batch accepts `/v1/responses` as an
input endpoint. That turns the decision into a control-plane choice over the
same underlying model surface. For teams building long-running AI jobs, that is
the distinction worth optimising.

My view is straightforward. Start with background mode when a single job still
belongs to your application flow, especially on agent backends where progress,
cancellation, and per-job debugging matter. Move to Batch when the work has
become a backlog problem and the 50 percent discount plus higher rate-limit pool
matter more than immediacy. That is a cleaner split than most launch summaries
give it.

## What does OpenAI background mode actually preserve?

Background mode preserves the normal Responses object as the unit your
application talks to. You create a request with `background: true`, receive a
response id back, then keep polling while status is `queued` or
`in_progress`. The same guide documents response-level cancellation, and it also
documents a hybrid path where you set both `background` and `stream` to true so
the job continues even if the client drops the connection. That is a real
control feature rather than a cosmetic wrapper around long latency.

The engineering implication is that your existing job model can often stay
intact. One request becomes one tracked object with one cancellation path and
one place to retrieve output. If you are already working with hosted OpenAI job
surfaces such as [the Agents API runtime boundary](/what-openai-agents-api-manages/),
background mode fits naturally because it behaves like a live job inside your
application. It feels very different from a later file import.

**Where the two async paths diverge**

1. Responses request with background true
   One job stays a normal response object.
2. Poll, stream, or cancel by response id
   Application keeps per-job control.
3. response object carries final output
   Read it back through Responses.
--- compare with Batch below ---
4. JSONL file of many requests
   Each line needs a custom_id.
5. 24 hour batch worker window
   Higher throughput, lower urgency.
6. output and error files
   Results arrive through files and `custom_id`.

Both paths can reach the same models and even the same `/v1/responses` endpoint. The real split is the control surface your application has to operate afterwards.

## What does the OpenAI Batch API add that background mode does not?

Batch adds three guarantees the background guide does not make: 50 percent lower
cost than synchronous APIs, a separate pool of significantly higher rate limits,
and a `24h` completion window. It also changes the input contract. Work starts
as a `.jsonl` file, each request line carries its own `custom_id`, and OpenAI
processes the set as one batch object rather than as individually retrievable
response ids. Even when the underlying request body targets `/v1/responses`, the
job-control layer is different.

That difference is why Batch suits offline operations better than interactive
ones. Evaluations, bulk classification, repository embeddings, and queued
content transformations usually want price and throughput more than they want a
stream that can resume mid-flight. The file boundary also makes ownership
cleaner for operations teams that already think in staged artifacts, similar to
the queue-first runtime pattern in [cloud agent deployment](/cloud-agents-manual-deployment/).
What you lose is the per-request feel of a live application call.

**Relative price index for the two async paths**

| Item | Value (index) | Note |
|---|---|---|
| Responses request, sync or background | 100 | baseline pricing path |
| Batch API request | 50 | guide states 50 percent lower cost |

Batch is cheaper by contract rather than by accident. If your job can tolerate the file workflow and the 24 hour window, cost pressure alone is a good reason to move it.

## How do OpenAI background mode and Batch differ on retention?

This is the detail most teams miss, and it is the strongest reason to read the
primary docs rather than secondary tutorials. OpenAI's background guide and
"Your data" guide both say background responses may be deleted after roughly 10
minutes when `store` is omitted or false, with a narrower rule for projects
using Modified Abuse Monitoring: long-term retention after the polling period
happens only when `store=true` is explicitly provided where policy permits it.
So background mode is operationally retrievable, but it is not a durable
results store by default.

Batch behaves differently. The batch object points to `output_file_id` and, if
needed, `error_file_id`, and the guide says the output file is automatically
deleted 30 days after the batch completes. That is still temporary, but it is a
very different operational envelope from roughly 10 minutes. If a downstream
system joins results slowly, Batch gives you much more slack. If your design
needs permanent history, neither path replaces writing the result into your own
database after retrieval.

## When should OpenAI background mode be your default?

Background mode should be your default when the important noun in the system is
"job", singular. A user submitted one research task. One internal service kicked
off one agent run. One operator may need to cancel one response when a budget
cap trips. In those situations, the response id is the right operational key,
and the ability to poll, cancel, or resume a stream is more valuable than batch
discounts. The control surface matches the mental model.

This is especially true for long-running agent work. An agent loop often wants
the same application to watch progress, tie logs to one request, and stop work
when a human intervenes. Batch can carry the same model request body, but the
cost of unpacking results from files usually exceeds the savings until volume is
high. My opinion is that teams reach for Batch too early because the 50 percent
number is loud. Control-plane friction is quieter, but it is what operators pay
for every day.

## When should the OpenAI Batch API win?

Batch should win when the important noun is "backlog". You have thousands of
requests, none of them need immediate delivery, and the surrounding system is
already comfortable with file upload, later collection, and result joins through
`custom_id`. That is the shape OpenAI describes directly in the guide: running
evals, classifying large datasets, and embedding content repositories. In those
cases, lower cost and more throughput are first-order advantages rather than
nice extras.

Batch is also the more honest fit when expiry is acceptable. The guide states
that batches can expire, with unfinished requests recorded in the error file,
and that any completed requests still appear in the output file. That partial
delivery model is normal for offline processing and awkward for user-facing job
control. If your consumer can reconcile a mixed output file and resubmit failed
or expired lines, Batch is doing exactly the job it was designed for. If not,
stay with background mode and keep the application loop simpler.

<ReadNext
  href="/cloud-agents-manual-deployment/"
  kicker="Related"
  title="How do cloud agents run in production?"
  note="The runtime controls that matter after you decide whether an AI job is interactive, background, or fully offline."
/>

## Choose between OpenAI background mode and the Batch API

Start with the application boundary, then move to urgency, price, and retention. That order removes the most expensive category errors first.

1. **Ask whether one job must stay inside your live application flow**: If another service, a user session, or an operator still needs one job id they can poll or cancel, start with background mode. It keeps the unit of work as a normal Responses object.
2. **Check whether the work can tolerate a file-based 24 hour lane**: Batch only offers a `24h` completion window and starts from a `.jsonl` upload. If that sounds normal for the job, such as nightly evals or mass classification, Batch is a candidate.
3. **Compare price pressure and throughput pressure together**: Batch wins when the backlog is large because OpenAI documents both a 50 percent discount and a separate pool of significantly higher rate limits. Background mode keeps ordinary request economics.
4. **Decide how your system will read results back**: Background responses are retrieved by response id and can also be streamed. Batch writes output and error files keyed by `custom_id`, so downstream joins and retries happen at the file layer.
5. **Review retention and deletion rules before launch**: Background mode may delete response data after roughly 10 minutes unless `store=true` is explicitly allowed and set. Batch output files persist much longer, but they are still temporary operational artifacts rather than a permanent database.
6. **Pilot one representative workload before standardising**: Run the same workload shape once through background mode and once through Batch. Measure operational friction alongside tokens, because the cheaper path can still be the wrong one if it breaks the surrounding system.


## Frequently asked questions

### Can the Batch API call the Responses API?

Yes. OpenAI's Batch guide lists `/v1/responses` as a supported endpoint. That means the background mode versus Batch decision is about delivery mechanics, rate limits, retention, and price rather than about a separate model feature set.

### Does background mode cost less than a normal Responses request?

The background guide does not advertise a discount. The Batch guide does advertise 50% lower cost than synchronous APIs, so cost-sensitive offline work belongs there unless you need the single-response control surface.

### Can I cancel both kinds of work?

Yes, but the cancellation unit is different. Background mode cancels one response object. Batch cancels the whole batch job, then waits for in-flight requests to finish for up to 10 minutes before the batch becomes cancelled.

### Which path is better for agent backends?

Background mode is usually better when a user or another service still cares about one job's status, partial progress, or cancellation. Batch is better when the work is a large offline set such as evaluations, repository embeddings, or bulk classification.

### What is the easiest mistake to make when switching to Batch?

Treating it like a cheaper background response. Batch changes the input shape to `.jsonl`, requires `custom_id` per request, and returns results through output files rather than a normal Responses retrieve call.


## Sources
- [Background mode](https://developers.openai.com/api/docs/guides/background). OpenAI, 2026-10-01
- [Batch API](https://developers.openai.com/api/docs/guides/batch). OpenAI, 2026-10-01
- [Your data](https://developers.openai.com/api/docs/guides/your-data). OpenAI, 2026-10-01

---
Canonical HTML: https://deepthinkingai.org/openai-background-mode-vs-batch-api/