Guides
Long-form notes on OpenAI service tiers, the 272K long-context cliff, prompt caching and how each lever actually multiplies a bill. Worked examples use the same price table as the calculator.
Tiers
- OpenAI Batch API: how the 50% discount works and when it is the wrong choice Batch is a 0.5× multiplier on the whole request, not a separate price list. A worked enrichment job, the 24-hour window that forbids agent loops, and the arithmetic of trimming first.
- OpenAI Fast mode: what 2x the rate actually buys you Fast is a 2x multiplier for up to 2.5x the speed, bought per request rather than per account. The clean result is that buying Fast for a share of traffic adds roughly that share to the bill.
- Flex vs Batch: both cost half, but they fail differently Two tiers halve the rate and differ only in how they break. Flex is synchronous and may be unavailable; Batch is asynchronous and bounded by 24 hours. A retry strategy for each.
- How to choose an OpenAI service tier for a request Batch and Flex both halve the rate, and they are the largest saving most teams never set. A decision order that maps nine common request types to a tier, a model and an effort level.
- OpenAI Ultrafast: 6x the price, 6x the speed, Astra only The 6x multiplier lands on input, cache read, cache write and output together, and only GPT-6 Astra publishes Ultrafast pricing. It pays only when wall-clock time is the binding constraint.