Eight levers that cut an OpenAI API bill, ranked by how much they save
A ranked list of the eight changes that actually move an OpenAI bill, from the half-price tiers down to the regional surcharge, with a quantified effect and a reason to skip each one.
Most API bills are not dominated by one mistake. They are dominated by a handful of defaults nobody revisited: everything on Standard, everything at the default effort, every prefix sent uncached. The eight levers below are ranked by how much they typically save, and each one comes with the case where you should leave it alone. The ranking assumes a mixed production bill; your mileage depends on which of these defaults you are still carrying.
The ranking
| # | lever | mechanism | typical saving | do not use when |
|---|---|---|---|---|
| 1 | Right tier per workload | 0.5× on non-interactive traffic | up to 50% of that traffic | a human is waiting |
| 2 | Low effort on Luna | output-side multiplier and a cheaper model | often an order of magnitude | the task needs reasoning |
| 3 | Cache the stable prefix | reads at 5–10% of input | 83% on a prefix reused 10× | the prefix changes every call |
| 4 | Stay under 272K | avoids a 1.98× re-pricing | about 50% per request | the whole document is required |
| 5 | Cap maxOutput | bounds the expensive column | worst-case only | output is already small |
| 6 | Stop defaulting to Fast | removes a 2× multiplier | the Fast share of traffic | a human is watching the request |
| 7 | Send spikes to Flex | 0.5× on retryable traffic | up to 50% of spike cost | requests cannot be retried |
| 8 | Audit regional endpoints | removes a 10% surcharge | 10% of that traffic | you have a residency requirement |
Lever 1: give each workload the right tier
The largest saving on the list is also the least glamorous. Batch and Flex are both 0.5× the Standard rate, and a large share of production traffic is neither interactive nor latency-sensitive. Moving it to a half-price tier cuts that traffic's cost in half with no change to the model, the prompt or the answer.
Start by sorting traffic into "someone is waiting" and "nobody is waiting". The second bucket goes to Batch if it is scheduled and Flex if it is retryable. Do not pull this lever for anything a person is watching, and do not set Flex account-wide — it turns your p99 latency into someone else's load.
| non-interactive share | saving on that share | effect on the total bill |
|---|---|---|
| 20% | 50% | −10% |
| 40% | 50% | −20% |
| 60% | 50% | −30% |
The table is the whole argument for auditing your traffic mix before anything else. A team with 60% non-interactive traffic has a 30% saving sitting in a parameter it has never set.
Lever 2: drop deterministic tasks to low effort on Luna
Classification, extraction, routing and formatting have one correct shape. Paying a flagship model to reason about them is waste, and the saving compounds because two things move at once: the model gets cheaper and the effort level drops.
GPT-6 Luna is $0.10 per million input and $0.50 per million output, against $2.00 and $10.00 on GPT-6.1 Sol. That is a 20× difference on input before effort is considered. Effort then scales the output side, and the low-to-high multiplier in the planning model is 8×. Together they can turn a noticeable line item into a rounding error.
The lever is wrong when the task actually needs reasoning. A classifier that mislabels edge cases because it was downgraded costs more in retries and corrections than it saves, so test the cheaper configuration against your own evaluation set before you ship it.
Lever 3: cache the stable prefix
Cached input is billed at a small fraction of the input rate — 5% on GPT-6.1 Sol, 10% on the other models — while a cache write costs 1.25×. A prefix reused ten times drops from $0.200 to $0.034, an 83% saving on that portion.
The lever fails when the prefix is not stable. If timestamps, request IDs or per-user data sit early in the prompt, every request writes a new cache entry and you pay the 1.25× premium for nothing. Move the static material first and the volatile material last.
Lever 4: keep input under 272K
Crossing 272,000 input tokens re-prices the whole request, roughly doubling it. On both Sol and Astra, one token over the line takes the bill to 1.98×. Trimming retrieval is usually the cheapest fix, followed by summarising history and chunking the job.
Do not pull this lever when the task genuinely needs the whole document. Whole-codebase refactors and legal review are cases where the multiplier is the honest price of the work. The mistake to avoid is crossing by accident rather than by choice.
Lever 5: cap maxOutput
Output is the expensive column — $10.00 per million on Sol against $2.00 for input, and $50.00 against $10.00 on Astra. A runaway generation is therefore the most expensive way to be wrong.
Setting a maxOutput cap does not reduce normal usage; it bounds the worst case. That is worth doing because it converts an unbounded risk into a known ceiling, and because a cap often surfaces bugs where a prompt is looping. It saves nothing on a workload whose output is already short, so treat it as insurance rather than an optimisation.
Lever 6: stop defaulting to Fast
Fast is 2× the Standard rate for up to 2.5× the speed, and it is priced per request. Doubling the rate on a share x of traffic adds about x to the bill, so enabling it account-wide doubles everything.
Buy it only for requests a human is watching. If you cannot name the route that needs it, you do not need it. Background work, scheduled jobs and anything already asynchronous belong on Standard or below.
Lever 7: send spikes to Flex
Flex is the right home for traffic you cannot forecast and can afford to retry. It halves the rate, and its failure mode — queueing or unavailability — is one a retry policy can absorb. Pair it with exponential backoff, a capped attempt count and a Standard fallback for requests that must answer now.
Do not use Flex for work that cannot be retried, and do not let it become the default. A request that falls back to Standard costs full price anyway, so an aggressive Flex rollout can add a second code path and save very little.
Lever 8: audit regional endpoints
Regional processing adds 10% on top of any rate, and it is only offered on Standard, Batch and Flex — Fast and Ultrafast do not support it. If you route traffic through a regional endpoint, confirm that every request on it actually needs to be there. Traffic that was pointed at a regional endpoint for one compliant workload and then left there for everything is the easiest 10% to recover.
Do not remove it if you have a genuine residency requirement. In that case the surcharge is the cost of compliance, and it belongs in your fixed assumptions rather than your waste column.
Putting it together
The levers are not equal, and they are not independent. Tier choice and effort move the most money, and they are the easiest to apply because they need no architectural change. Caching and the cliff need you to understand the prompt, and that is where the biggest surprises live. The rest — maxOutput, Fast, Flex, regional — are guards against specific mistakes rather than general savings.
A sensible order of operations: pick the tier, then the effort, then cache the prefix, then check the input count, then cap the output, then audit the latency and residency choices. Done in that order, each change is measurable on its own, and you can attribute the saving to the lever that produced it.
The short version
- The biggest saving is the least glamorous: half-price tiers for non-interactive traffic.
- Dropping deterministic work to low effort on Luna compounds a cheaper model with a smaller output budget.
- Caching cuts a reused prefix by up to 94%; it needs a stable prefix to work.
- Staying under 272K avoids a 1.98× re-pricing, and trimming retrieval is usually free.
- maxOutput caps bound the worst case; they are insurance, not a saving.
- Fast doubles the rate, Flex halves it, and the regional surcharge adds 10% — audit all three rather than assuming the defaults.
Frequently asked
What is the single largest saving available?
Moving non-interactive traffic to Batch or Flex, both 0.5x the Standard rate. It is the largest saving and the one most accounts never set.
Should I always use the cheapest model?
No. The cheapest model is right for deterministic tasks like classification and routing, but using it for reasoning-heavy work tends to cost more once retries and corrections are counted.
Does capping maxOutput save money?
It bounds the worst case rather than reducing the typical case. It is worth setting because output is the expensive column, but it will not shrink a bill where output is already small.
Is the regional 10% surcharge avoidable?
Only if you do not have a residency requirement. If you do, the surcharge is the cost of compliance and should be treated as fixed, not as waste.