Tiers

Flex vs Batch: both cost half, but they fail differently

Two tiers halve the rate and differ only in how they break. Flex is synchronous and may be unavailable; Batch is asynchronous and bounded by 24 hours. A retry strategy for each.

Batch and Flex are both billed at 0.5× the Standard rate, so price does not separate them. Their failure modes do. Batch fails by taking up to 24 hours; Flex fails by queueing, or by being unavailable exactly when you need it. Choosing between them is choosing which failure your system can absorb, and that is an engineering question rather than a pricing one.

Getting this wrong is expensive in a subtle way. Teams pick Flex because it looks like Standard with a discount, discover the p99 latency belongs to someone else's load, add a Standard fallback, and end up paying full price for the traffic anyway — having taken on the complexity for nothing.

The only difference that matters

BatchFlex
multiplier0.5×0.5×
modeasynchronous, file-basedsynchronous
waitup to 24 hours, boundedqueued, unbounded
under loadqueuedmay be unavailable
regional processingyesyes
retrypart of the job modelyour responsibility

Both halve the rate. Both support regional processing, which adds 10% on top. Everything that actually differs is in the middle four rows, and all of it is about time and availability.

How Flex fails

Flex is a lower-priority synchronous tier. It answers in the same request cycle as Standard when capacity is available, and it may queue or return unavailable when it is not. There is no deadline attached, so you cannot reason about the worst case the way you can with Batch.

That makes Flex a good fit for work that is genuinely optional in the moment: a non-critical UI hint, a background suggestion, a request whose caller is willing to retry. It is a bad fit for anything a user is watching, because their latency is now a function of demand you cannot see.

How Batch fails

Batch fails differently and more predictably. It is asynchronous, so there is no latency to speak of — only a completion window of up to 24 hours. If the result is not ready, you have not been surprised; you signed up for it.

The failure mode is therefore "late", not "unavailable", and that is a much easier failure to design around. A job that runs on a schedule already tolerates being late. A job that answers a live request does not.

A retry strategy for Flex

If you route retryable traffic to Flex, the retry policy is not optional — it is the whole safety mechanism.

A half-price tier that you retry on Standard is not a discount. It is Standard with extra steps, plus a failure mode you now own.

Why Flex as a default bites

The tempting move is to set Flex account-wide and bank the 50%. It usually backfires for three reasons.

First, it makes your p99 latency a function of external load, which is the opposite of what most products want. Second, it turns a deterministic system into a retrying one, and retries hide bugs until they do not. Third, the requests that do fall back to Standard cost full price anyway, so the blended saving is much smaller than 50% and often not worth the added branch.

Flex works best as a targeted route: a named slice of traffic that is retryable, non-critical, and cheap to lose. It works worst as a global setting.

Side by side, by workload

workloadBatchFlexwhy
Overnight enrichmentyesnoNobody is waiting, so take the bounded queue.
Evaluation runyesnoFile-based by nature; no need for synchronous responses.
Unpredictable traffic spikenoyesSynchronous and retryable, but no fixed deadline to work against.
Non-critical UI suggestionnoyesThe product tolerates a miss, and the caller can retry.
Interactive chatnonoA human is waiting; both half-price tiers are disqualified.
Agent loopnonoNeeds each tool result inside the same request cycle.

The pattern is that Batch wants work with a schedule, and Flex wants work with a tolerance for missing. Anything with a person attached to it wants neither, and the safest default when you are unsure is to leave it on Standard until you can measure it.

Deciding with a number

The choice stops being a debate once you put a number on it. Estimate the share of traffic that is scheduled and non-interactive, and the share that is interactive but retryable. Those two shares are your Batch and Flex budgets.

Take a million requests a month, 60% of them background work that can be scheduled and 15% of them retryable-but-interactive. Moving the scheduled slice to Batch halves 60% of the bill, a 30% saving on the total. Moving the retryable slice to Flex halves another 15%, for a further 7.5%. The remaining 25% stays on Standard because a human is waiting.

The estimate is only as good as the fallback rate. If a third of the Flex traffic ends up retried on Standard, the blended saving on that slice falls from 50% to roughly 33%, and the extra code path starts to look expensive for the money it returns.

Running both at once

Batch and Flex are not alternatives at the account level; they are routes. A healthy setup uses both, with interactive traffic left alone.

traffictiersignal
Nightly enrichmentBatchscheduled, nobody watching
Evaluation runsBatchfile-based and rerunnable
Backfills and cache warm-upsBatchno deadline
Retryable UI hintsFlexcan miss, caller retries
Unforecastable spikesFlexnon-critical, tolerant of queueing
Live chat and agent loopsStandarda human is waiting

The table is a routing policy, not a purchase. Nothing in it is an account-wide setting, which is what makes it safe to adopt incrementally.

What to instrument

Both half-price tiers need the same two measurements: the share of traffic routed to them, and the rate at which that traffic fails to complete on the cheap path.

For Flex, track the unavailable rate and the fallback-to-Standard rate. If the second is high, Flex is not saving money — it is a Standard request with extra steps. For Batch, track the completion-time distribution and the partial-failure rate, because a job that finishes within 24 hours but loses 5% of its requests is not half price once you rerun the failures.

Without those two numbers you are guessing, and the guess usually flatters the tier. Instrument first, then route.

The short version

Frequently asked

Is Flex cheaper than Batch?

No. Both are billed at 0.5× the Standard rate. The difference is scheduling and failure mode, not price.

What happens when Flex is unavailable?

The request may queue or fail rather than complete immediately. There is no bounded deadline, so a caller that cannot wait must retry, fall back to Standard, or drop the work.

Does Batch or Flex support regional processing?

Both do. Regional processing endpoints add a further 10% on top of the rate, and both half-price tiers support them.

Which should I use for a traffic spike?

Flex, if the requests are retryable and the product tolerates queueing. Batch is wrong for a spike because it is asynchronous and has no way to answer in the same request cycle.