Cost

Eight levers that cut an OpenAI API bill, ranked by how much they save

A ranked list of the eight changes that actually move an OpenAI bill, from the half-price tiers down to the regional surcharge, with a quantified effect and a reason to skip each one.

Most API bills are not dominated by one mistake. They are dominated by a handful of defaults nobody revisited: everything on Standard, everything at the default effort, every prefix sent uncached. The eight levers below are ranked by how much they typically save, and each one comes with the case where you should leave it alone. The ranking assumes a mixed production bill; your mileage depends on which of these defaults you are still carrying.

The ranking

#levermechanismtypical savingdo not use when
1Right tier per workload0.5× on non-interactive trafficup to 50% of that traffica human is waiting
2Low effort on Lunaoutput-side multiplier and a cheaper modeloften an order of magnitudethe task needs reasoning
3Cache the stable prefixreads at 5–10% of input83% on a prefix reused 10×the prefix changes every call
4Stay under 272Kavoids a 1.98× re-pricingabout 50% per requestthe whole document is required
5Cap maxOutputbounds the expensive columnworst-case onlyoutput is already small
6Stop defaulting to Fastremoves a 2× multiplierthe Fast share of traffica human is watching the request
7Send spikes to Flex0.5× on retryable trafficup to 50% of spike costrequests cannot be retried
8Audit regional endpointsremoves a 10% surcharge10% of that trafficyou have a residency requirement

Lever 1: give each workload the right tier

The largest saving on the list is also the least glamorous. Batch and Flex are both 0.5× the Standard rate, and a large share of production traffic is neither interactive nor latency-sensitive. Moving it to a half-price tier cuts that traffic's cost in half with no change to the model, the prompt or the answer.

Start by sorting traffic into "someone is waiting" and "nobody is waiting". The second bucket goes to Batch if it is scheduled and Flex if it is retryable. Do not pull this lever for anything a person is watching, and do not set Flex account-wide — it turns your p99 latency into someone else's load.

non-interactive sharesaving on that shareeffect on the total bill
20%50%−10%
40%50%−20%
60%50%−30%

The table is the whole argument for auditing your traffic mix before anything else. A team with 60% non-interactive traffic has a 30% saving sitting in a parameter it has never set.

Lever 2: drop deterministic tasks to low effort on Luna

Classification, extraction, routing and formatting have one correct shape. Paying a flagship model to reason about them is waste, and the saving compounds because two things move at once: the model gets cheaper and the effort level drops.

GPT-6 Luna is $0.10 per million input and $0.50 per million output, against $2.00 and $10.00 on GPT-6.1 Sol. That is a 20× difference on input before effort is considered. Effort then scales the output side, and the low-to-high multiplier in the planning model is 8×. Together they can turn a noticeable line item into a rounding error.

The lever is wrong when the task actually needs reasoning. A classifier that mislabels edge cases because it was downgraded costs more in retries and corrections than it saves, so test the cheaper configuration against your own evaluation set before you ship it.

Lever 3: cache the stable prefix

Cached input is billed at a small fraction of the input rate — 5% on GPT-6.1 Sol, 10% on the other models — while a cache write costs 1.25×. A prefix reused ten times drops from $0.200 to $0.034, an 83% saving on that portion.

The lever fails when the prefix is not stable. If timestamps, request IDs or per-user data sit early in the prompt, every request writes a new cache entry and you pay the 1.25× premium for nothing. Move the static material first and the volatile material last.

Lever 4: keep input under 272K

Crossing 272,000 input tokens re-prices the whole request, roughly doubling it. On both Sol and Astra, one token over the line takes the bill to 1.98×. Trimming retrieval is usually the cheapest fix, followed by summarising history and chunking the job.

Do not pull this lever when the task genuinely needs the whole document. Whole-codebase refactors and legal review are cases where the multiplier is the honest price of the work. The mistake to avoid is crossing by accident rather than by choice.

Lever 5: cap maxOutput

Output is the expensive column — $10.00 per million on Sol against $2.00 for input, and $50.00 against $10.00 on Astra. A runaway generation is therefore the most expensive way to be wrong.

Setting a maxOutput cap does not reduce normal usage; it bounds the worst case. That is worth doing because it converts an unbounded risk into a known ceiling, and because a cap often surfaces bugs where a prompt is looping. It saves nothing on a workload whose output is already short, so treat it as insurance rather than an optimisation.

Lever 6: stop defaulting to Fast

Fast is 2× the Standard rate for up to 2.5× the speed, and it is priced per request. Doubling the rate on a share x of traffic adds about x to the bill, so enabling it account-wide doubles everything.

Buy it only for requests a human is watching. If you cannot name the route that needs it, you do not need it. Background work, scheduled jobs and anything already asynchronous belong on Standard or below.

Lever 7: send spikes to Flex

Flex is the right home for traffic you cannot forecast and can afford to retry. It halves the rate, and its failure mode — queueing or unavailability — is one a retry policy can absorb. Pair it with exponential backoff, a capped attempt count and a Standard fallback for requests that must answer now.

Do not use Flex for work that cannot be retried, and do not let it become the default. A request that falls back to Standard costs full price anyway, so an aggressive Flex rollout can add a second code path and save very little.

Lever 8: audit regional endpoints

Regional processing adds 10% on top of any rate, and it is only offered on Standard, Batch and Flex — Fast and Ultrafast do not support it. If you route traffic through a regional endpoint, confirm that every request on it actually needs to be there. Traffic that was pointed at a regional endpoint for one compliant workload and then left there for everything is the easiest 10% to recover.

Do not remove it if you have a genuine residency requirement. In that case the surcharge is the cost of compliance, and it belongs in your fixed assumptions rather than your waste column.

Putting it together

The levers are not equal, and they are not independent. Tier choice and effort move the most money, and they are the easiest to apply because they need no architectural change. Caching and the cliff need you to understand the prompt, and that is where the biggest surprises live. The rest — maxOutput, Fast, Flex, regional — are guards against specific mistakes rather than general savings.

A sensible order of operations: pick the tier, then the effort, then cache the prefix, then check the input count, then cap the output, then audit the latency and residency choices. Done in that order, each change is measurable on its own, and you can attribute the saving to the lever that produced it.

The short version

Frequently asked

What is the single largest saving available?

Moving non-interactive traffic to Batch or Flex, both 0.5x the Standard rate. It is the largest saving and the one most accounts never set.

Should I always use the cheapest model?

No. The cheapest model is right for deterministic tasks like classification and routing, but using it for reasoning-heavy work tends to cost more once retries and corrections are counted.

Does capping maxOutput save money?

It bounds the worst case rather than reducing the typical case. It is worth setting because output is the expensive column, but it will not shrink a bill where output is already small.

Is the regional 10% surcharge avoidable?

Only if you do not have a residency requirement. If you do, the surcharge is the cost of compliance and should be treated as fixed, not as waste.