Tiers

How to choose an OpenAI service tier for a request

Batch and Flex both halve the rate, and they are the largest saving most teams never set. A decision order that maps nine common request types to a tier, a model and an effort level.

Most teams pay standard list price for everything, because choosing a tier feels like a performance decision. It is a queueing decision. Three questions settle it, and you should ask them in order: is a human waiting on this response, can the job be retried without consequence, and is latency part of what the product sells? The first yes decides the tier. Everything after that is a model and an effort choice.

This matters because two of the five tiers cost half as much as the default, and most accounts never set either of them. That is the largest single saving available to a normal API bill, and it is not a pricing trick — it is a scheduling trade you are allowed to take.

Ask who is waiting

The first question removes most of the ambiguity. A person watching a response arrive will notice a queue. A nightly job will not.

A surprising amount of traffic that "feels" interactive is not. A nightly enrichment pass, an evaluation run, an embedding refresh, a moderation sweep and a document backfill are all Batch workloads that teams leave on Standard because the code was written as a loop that calls the API synchronously and waits.

Ask whether the job can retry

The second question decides between Batch and Flex, which cost the same.

Batch is the easier trade to reason about: half price, asynchronous, and a documented completion window of up to 24 hours. Flex is also half price, but it is a lower-priority synchronous tier — it may queue, and it may be unavailable under load.

If a workload can be retried safely and you do not control when it runs, Flex is a reasonable home for it. If it genuinely cannot fail silently and nobody is watching, Batch's bounded 24-hour window is easier to defend than a tier that might simply return unavailable at 3am.

Both Batch and Flex are billed at 0.5× the Standard rate. The difference is not price — it is whether you are buying a queue with a deadline, or a lower priority with no deadline at all.

Ask whether latency is the product

Only after the first two answers come back "someone is waiting" and "it cannot queue" should you consider paying above Standard.

Fast is 2× the Standard rate for up to 2.5× the speed. It was called priority until 2026-07-30, and it is bought per request, not per account. Ultrafast is 6× the rate for up to 6× the speed, is published only for GPT-6 Astra, and is not available on any other model.

Both are for requests where wall-clock time is a constraint the product feels. Neither supports regional processing endpoints, so a residency requirement rules them out. If you have both a residency and a latency requirement, you cannot buy your way out of it with a tier — you have to fix the latency somewhere else.

Nine request types, mapped

The table below is the same opinion layer the calculator ships with. It maps a request type to a tier, a model and an effort level.

request typetiermodeleffortwhy
Overnight enrichment or backfillBatchGPT-6.1 SollowNobody is waiting. Half price is the whole decision.
High-volume classification or routingBatchGPT-6 LunalowA label has one correct shape. Reasoning cannot improve a lookup.
Customer-facing chatStandardGPT-6.1 SolmediumA human is waiting and traffic is unpredictable. Cache the system prompt instead of buying speed.
RAG over long documentsStandardGPT-6.1 SolmediumRetrieved context is your input line. Trim below 272K before considering a bigger tier.
Agent loop with many tool callsStandardGPT-6.1 SolhighHistory grows every turn — watch the 272K cliff and cache the stable prefix.
Traffic spike you cannot forecastFlexGPT-6.1 SolmediumFlex halves the rate and accepts queueing. Keep Standard for the requests that must answer now.
Latency-critical interactive UIFastGPT-6 Astrahigh2× standard. Buy it per request, never as an account-wide default.
Bulk generation, wall-clock boundUltrafastGPT-6 Astrahigh6× the price for up to 6× the speed. Astra only, and no EU endpoint.
One-off hard problemStandardGPT-6 AstraxhighEffort, not tier, is the right lever when a single answer is expensive to get wrong.

Two patterns fall out of that table. Five of the nine rows are Batch or Flex, so the half-price tiers are the correct answer far more often than their adoption suggests. And the premium tiers appear once each, both only when latency is the product rather than an implementation detail.

The multipliers, in one table

Every tier is a multiplier on the model's Standard rate. That is the entire mechanism, and it is worth memorising because it composes with everything else.

tiermultiplierspeed ceilingregional processing
Batch0.5×up to 24 hoursyes
Flex0.5×queued, may be unavailableyes
Standard1×baselineyes
Fast2×up to 2.5×no
Ultrafast6×up to 6×no

Read the last column carefully. Regional processing adds 10% on top of whatever rate you are paying, and the two fast tiers do not offer it at all. A data-residency requirement and a latency requirement cannot both be satisfied by the tier ladder alone.

Where effort fits

Tier is about scheduling. Effort is about how much work the model does, and it moves the output side of the bill only. Input tokens cost the same at every effort level.

The two levers therefore act on different halves of the cost equation. Batch and Flex are multipliers on the whole request, so they touch input and output together. Effort touches output alone.

If your bill is dominated by input — which it is for RAG and long agent loops — effort is the wrong lever. Fix the input instead: trim retrieval, summarise history, cache the stable prefix, and stay under the 272K cliff. If your bill is dominated by output, effort is the right lever, and moving a deterministic task from high down to low is often a larger saving than any tier change.

The effort multipliers used throughout this site — 1, 3, 8, 20 and 45 for low through max — are a planning assumption, not a published vendor figure. Use them to compare options, then replace them with your own measurements.

What to change first

If you change one thing this week, move your non-interactive traffic off Standard. The second-largest lever is effort on deterministic tasks. The rest — Fast, Ultrafast, regional endpoints — are multipliers to reach for only when a specific request genuinely needs them.

The short version

Frequently asked

What is the default service tier?

Standard. It is synchronous, billed at list price, and applies whenever you do not set a tier explicitly on the request.

Can the tier be changed per request?

Yes. Tiers are chosen per request rather than per account, so the same model can serve interactive traffic at Standard and background traffic at Batch or Flex.

Which tier is cheapest?

Batch and Flex, both at half the Standard rate. They differ in failure mode rather than price: Batch accepts a wait of up to 24 hours, Flex may queue or be unavailable.

Does the tier change the quality of the answer?

No. The tier changes how the request is scheduled and what it costs, not which model answers or how hard it reasons. Effort is the lever for answer quality.