Flex vs Batch: both cost half, but they fail differently
Two tiers halve the rate and differ only in how they break. Flex is synchronous and may be unavailable; Batch is asynchronous and bounded by 24 hours. A retry strategy for each.
Batch and Flex are both billed at 0.5× the Standard rate, so price does not separate them. Their failure modes do. Batch fails by taking up to 24 hours; Flex fails by queueing, or by being unavailable exactly when you need it. Choosing between them is choosing which failure your system can absorb, and that is an engineering question rather than a pricing one.
Getting this wrong is expensive in a subtle way. Teams pick Flex because it looks like Standard with a discount, discover the p99 latency belongs to someone else's load, add a Standard fallback, and end up paying full price for the traffic anyway — having taken on the complexity for nothing.
The only difference that matters
| Batch | Flex | |
|---|---|---|
| multiplier | 0.5× | 0.5× |
| mode | asynchronous, file-based | synchronous |
| wait | up to 24 hours, bounded | queued, unbounded |
| under load | queued | may be unavailable |
| regional processing | yes | yes |
| retry | part of the job model | your responsibility |
Both halve the rate. Both support regional processing, which adds 10% on top. Everything that actually differs is in the middle four rows, and all of it is about time and availability.
How Flex fails
Flex is a lower-priority synchronous tier. It answers in the same request cycle as Standard when capacity is available, and it may queue or return unavailable when it is not. There is no deadline attached, so you cannot reason about the worst case the way you can with Batch.
That makes Flex a good fit for work that is genuinely optional in the moment: a non-critical UI hint, a background suggestion, a request whose caller is willing to retry. It is a bad fit for anything a user is watching, because their latency is now a function of demand you cannot see.
How Batch fails
Batch fails differently and more predictably. It is asynchronous, so there is no latency to speak of — only a completion window of up to 24 hours. If the result is not ready, you have not been surprised; you signed up for it.
The failure mode is therefore "late", not "unavailable", and that is a much easier failure to design around. A job that runs on a schedule already tolerates being late. A job that answers a live request does not.
A retry strategy for Flex
If you route retryable traffic to Flex, the retry policy is not optional — it is the whole safety mechanism.
- Only send work that can be retried without side effects. A duplicate write is not retryable work.
- Use exponential backoff with jitter, not a tight loop, so a busy moment does not become a stampede.
- Cap the number of attempts, then fall back to Standard for that request.
- Keep a Standard path for anything that must answer now, and route by policy rather than by hope.
- Log the unavailable rate. If it is high, Flex is not saving money — it is adding a second code path that costs full price half the time.
A half-price tier that you retry on Standard is not a discount. It is Standard with extra steps, plus a failure mode you now own.
Why Flex as a default bites
The tempting move is to set Flex account-wide and bank the 50%. It usually backfires for three reasons.
First, it makes your p99 latency a function of external load, which is the opposite of what most products want. Second, it turns a deterministic system into a retrying one, and retries hide bugs until they do not. Third, the requests that do fall back to Standard cost full price anyway, so the blended saving is much smaller than 50% and often not worth the added branch.
Flex works best as a targeted route: a named slice of traffic that is retryable, non-critical, and cheap to lose. It works worst as a global setting.
Side by side, by workload
| workload | Batch | Flex | why |
|---|---|---|---|
| Overnight enrichment | yes | no | Nobody is waiting, so take the bounded queue. |
| Evaluation run | yes | no | File-based by nature; no need for synchronous responses. |
| Unpredictable traffic spike | no | yes | Synchronous and retryable, but no fixed deadline to work against. |
| Non-critical UI suggestion | no | yes | The product tolerates a miss, and the caller can retry. |
| Interactive chat | no | no | A human is waiting; both half-price tiers are disqualified. |
| Agent loop | no | no | Needs each tool result inside the same request cycle. |
The pattern is that Batch wants work with a schedule, and Flex wants work with a tolerance for missing. Anything with a person attached to it wants neither, and the safest default when you are unsure is to leave it on Standard until you can measure it.
Deciding with a number
The choice stops being a debate once you put a number on it. Estimate the share of traffic that is scheduled and non-interactive, and the share that is interactive but retryable. Those two shares are your Batch and Flex budgets.
Take a million requests a month, 60% of them background work that can be scheduled and 15% of them retryable-but-interactive. Moving the scheduled slice to Batch halves 60% of the bill, a 30% saving on the total. Moving the retryable slice to Flex halves another 15%, for a further 7.5%. The remaining 25% stays on Standard because a human is waiting.
The estimate is only as good as the fallback rate. If a third of the Flex traffic ends up retried on Standard, the blended saving on that slice falls from 50% to roughly 33%, and the extra code path starts to look expensive for the money it returns.
Running both at once
Batch and Flex are not alternatives at the account level; they are routes. A healthy setup uses both, with interactive traffic left alone.
| traffic | tier | signal |
|---|---|---|
| Nightly enrichment | Batch | scheduled, nobody watching |
| Evaluation runs | Batch | file-based and rerunnable |
| Backfills and cache warm-ups | Batch | no deadline |
| Retryable UI hints | Flex | can miss, caller retries |
| Unforecastable spikes | Flex | non-critical, tolerant of queueing |
| Live chat and agent loops | Standard | a human is waiting |
The table is a routing policy, not a purchase. Nothing in it is an account-wide setting, which is what makes it safe to adopt incrementally.
What to instrument
Both half-price tiers need the same two measurements: the share of traffic routed to them, and the rate at which that traffic fails to complete on the cheap path.
For Flex, track the unavailable rate and the fallback-to-Standard rate. If the second is high, Flex is not saving money — it is a Standard request with extra steps. For Batch, track the completion-time distribution and the partial-failure rate, because a job that finishes within 24 hours but loses 5% of its requests is not half price once you rerun the failures.
Without those two numbers you are guessing, and the guess usually flatters the tier. Instrument first, then route.
The short version
- Batch and Flex are both 0.5×; price does not choose between them.
- Batch fails by being late, bounded at 24 hours. Flex fails by queueing or being unavailable, unbounded.
- Flex needs a retry policy — backoff, a capped attempt count, and a Standard fallback.
- Setting Flex account-wide makes your p99 someone else's load and rarely saves what it promises.
- Instrument the fallback rate on Flex and the partial-failure rate on Batch before trusting either.
- Anything a human is waiting on belongs on Standard, not on either half-price tier.
Frequently asked
Is Flex cheaper than Batch?
No. Both are billed at 0.5× the Standard rate. The difference is scheduling and failure mode, not price.
What happens when Flex is unavailable?
The request may queue or fail rather than complete immediately. There is no bounded deadline, so a caller that cannot wait must retry, fall back to Standard, or drop the work.
Does Batch or Flex support regional processing?
Both do. Regional processing endpoints add a further 10% on top of the rate, and both half-price tiers support them.
Which should I use for a traffic spike?
Flex, if the requests are retryable and the product tolerates queueing. Batch is wrong for a spike because it is asynchronous and has no way to answer in the same request cycle.