OpenAI Batch API: how the 50% discount works and when it is the wrong choice
Batch is a 0.5× multiplier on the whole request, not a separate price list. A worked enrichment job, the 24-hour window that forbids agent loops, and the arithmetic of trimming first.
Batch is the simplest discount on the OpenAI price list: 0.5× the Standard rate, in exchange for patience. If a job can wait up to 24 hours, it costs half. If anything or anyone is waiting on the result, it does not belong here — and no amount of careful arithmetic will change that.
The teams that get the most out of Batch are not the ones with the cleverest prompts. They are the ones who noticed that a large slice of their traffic was never interactive in the first place.
What the 50% actually is
Batch is a multiplier, not a separate price list. The model rates stay the same; the whole request is billed at half. On GPT-6.1 Sol, an input token that costs $2.00 per million at Standard costs $1.00 per million in Batch, and output falls from $10.00 to $5.00 per million.
| tier | input (GPT-6.1 Sol) | output (GPT-6.1 Sol) | relative |
|---|---|---|---|
| Standard | $2.00 | $10.00 | 1× |
| Batch | $1.00 | $5.00 | 0.5× |
| Flex | $1.00 | $5.00 | 0.5× |
| Fast | $4.00 | $20.00 | 2× |
Rates are USD per million tokens. Every other multiplier still applies on top, including the long-context cliff and any regional processing surcharge. Batch halves the number; it does not exempt you from the others.
That composition is the reason Batch is worth understanding rather than just switching on. It is not a separate product with separate rules. It is a 0.5× you can place anywhere in a chain of multipliers, and where you place it relative to the others changes the answer.
The 24-hour window, and what it forbids
Batch is asynchronous and file-based. You submit a set of requests, and you collect the results later; the service targets completion within 24 hours. That window is what you are selling in exchange for the discount.
It rules out anything that needs an answer inside the current request cycle:
- Interactive chat, autocomplete, or any surface with a spinner
- Agent loops that call tools and need each result to decide the next step
- Streaming responses of any kind
- Moderation or validation that gates an action in the same transaction
If a workflow can be restructured so the model call happens before or after the user-facing transaction rather than inside it, Batch becomes available. That restructuring is usually the real work, and it is usually worth it.
Workloads that fit Batch
The pattern is "nobody is watching, and a rerun is free". Anything that matches gets an immediate 50% cut.
- Overnight enrichment and backfills of historical records
- Evaluation and regression runs across a fixed prompt set
- Embedding refreshes over a document corpus
- Bulk classification, tagging and moderation sweeps
- Synthetic data generation and large translation passes
- Any loop where the surrounding code already handles retries and partial failure
Workloads that do not fit
The mirror image is "something downstream is blocked". That includes interactive surfaces, but also subtler cases: a nightly report that must be ready before the morning stand-up, a same-day customer email, a batch job that itself feeds a live dashboard. The 24-hour window is an upper bound, not a promise, so anything with a deadline tighter than a day should not depend on it.
A worked enrichment job
Take a corpus of one million documents a month. Each request sends 2,500 input tokens and asks for a 300-token structured summary, on GPT-6.1 Sol.
| per call | per month | |
|---|---|---|
| Standard | $0.0080 | $8,000 |
| Batch | $0.0040 | $4,000 |
The saving is $4,000 a month, or $48,000 a year, from changing one parameter. No prompt engineering, no model downgrade, no quality trade — the same model produces the same summary. The only thing you give up is the guarantee that the answer arrives within the same second you asked for it, which for a nightly enrichment pass was never a guarantee you had.
Batch and the 272K cliff together
The two multipliers compose, and the order in which you apply them matters more than people expect. Consider a request with 300,000 input tokens and 2,000 output tokens on GPT-6.1 Sol — just over the 272,000-token long-context threshold, so the whole request is re-priced at input ×2 and output ×1.5.
| scenario | input cost | output cost | per call |
|---|---|---|---|
| Standard, over the cliff | $1.200 | $0.030 | $1.230 |
| Batch, still over the cliff | $0.600 | $0.015 | $0.615 |
| Trimmed to 270K, then Batch | $0.540 | $0.015 | $0.285 |
Trimming below the threshold before batching is 2.16× cheaper than batching alone. The reason is that Batch halves an already-inflated number, while trimming removes the inflation first and then halves what is left. If you have a long-context Batch job, look at the input count before you celebrate the discount.
Structuring a Batch job
Batch is file-based, so the shape of the job matters as much as the shape of a single call.
- Group requests that share a prompt, so the work is easy to rerun and easy to reason about.
- Expect partial failures. A batch returns per-request results and some will error, so your code must handle a mixture rather than assuming all-or-nothing.
- Keep the job idempotent. Because you will rerun it, a duplicate write is the most common Batch bug.
- Log input and output token counts per request. The discount is easy to verify and easy to lose if the job silently starts sending more context.
The engineering cost is real, which is why Batch pays off at volume and not at toy scale. A pipeline that submits, waits and reconciles results is worth building once and reusing, not rebuilding for every job.
Batch composes with the other levers
Batch is a multiplier, so it stacks with everything else on the price list. Caching still applies, and cache reads and writes are both halved along with the rest of the request, so a stable prefix is cheaper in Batch than at Standard. Effort still applies too: moving a Batch job from high to low effort cuts the output side before the 0.5× lands on top of it.
Batch halves every figure on the request, including the cached ones. It does not exempt you from the long-context multiplier, so a badly-trimmed long request is still a badly-trimmed long request after the discount.
The one lever that interacts awkwardly is the cliff, and the order in which you apply the two changes decides how much you save.
When Batch is the wrong choice
The discount is real but not free, and there are cases where taking it costs more than it saves.
- The job has a deadline tighter than 24 hours, whether that is an SLA or a colleague waiting for a file.
- The job is small. A file-based submission pipeline has real engineering cost, and it is not worth building for a few dollars a month.
- You are still iterating on the prompt. Each iteration costs a 24-hour round trip, which slows development far more than the discount is worth.
- The workload is multi-turn or agentic, so the 24-hour window breaks the control flow.
- The output feeds something that must react in real time.
The honest test is simple. If you can rerun the job tomorrow with no consequence, Batch is correct. If a rerun would embarrass someone, it is not.
The short version
- Batch is a 0.5× multiplier on the whole request, not a separate price list.
- The 24-hour window is the product; anything waiting on the result is disqualified.
- Overnight enrichment, evals, embedding refreshes and bulk classification all fit.
- Agent loops, streaming and same-transaction moderation do not.
- Batch and the 272K cliff compose: trim below the threshold first, then batch, and the saving compounds.
- If a rerun would be embarrassing, the discount is not yours to take.
Frequently asked
How much does the Batch API cost?
Half the Standard rate. Batch is a 0.5× multiplier on the whole request, so on GPT-6.1 Sol input drops from $2.00 to $1.00 per million tokens and output from $10.00 to $5.00 per million.
How long does Batch take?
The service targets completion within 24 hours. That window is the product you are buying with the discount, so the discount is not available to anything that needs an answer sooner.
Can I use Batch for an agent loop?
No. Batch is asynchronous and file-based, so it cannot feed a tool result back into the next step of the same conversation. Agent loops belong on Standard.
Does Batch support regional processing?
Yes. Unlike Fast and Ultrafast, Batch supports regional processing endpoints, which add a further 10% on top of the rate.