OpenAI Ultrafast: 6x the price, 6x the speed, Astra only
The 6x multiplier lands on input, cache read, cache write and output together, and only GPT-6 Astra publishes Ultrafast pricing. It pays only when wall-clock time is the binding constraint.
Ultrafast is 6× the Standard rate for up to 6× the speed, and only GPT-6 Astra has published pricing for it. It is the most expensive line on the price list and the easiest to waste money on, because the multiplier lands on all four figures at once rather than on output alone. This is a tier you buy to hit a deadline, not one you buy to save money.
The useful way to think about it is as a wall-clock purchase. Ultrafast does not make the model better, and it does not reduce the work. It compresses the time the work takes, and you pay six times the standard rate for that compression.
What 6× multiplies
On GPT-6 Astra, every rate is multiplied by six. That includes the cached-input rate, which is easy to overlook because it is small — until it is multiplied.
| figure | Standard | Ultrafast |
|---|---|---|
| input | $10.00 | $60.00 |
| cache read | $1.00 | $6.00 |
| cache write | $12.50 | $75.00 |
| output | $50.00 | $300.00 |
Rates are USD per million tokens. Note that caching still helps proportionally — a cached read at $6.00 is still a tenth of an uncached input at $60.00 — but the absolute numbers are large enough that a cache miss on an Ultrafast workload is expensive.
Only Astra has it
GPT-6 Astra is the only model with a published Ultrafast price. GPT-6.1 Sol Ultrafast has been announced but is not yet open, and the other models on the price list do not offer the tier at all. If you are not on Astra, the Ultrafast decision does not exist for you yet.
That narrow availability matters for planning. Any comparison of "should we use Ultrafast" is also a comparison of "should this workload run on Astra", because the two are inseparable. A team that cannot move to Astra does not have an Ultrafast option at all, and its choice collapses back to Standard against Fast.
A bulk run, priced
Take a bulk generation job of 10,000 requests, each sending 4,000 input tokens and producing 2,000 output tokens, on GPT-6 Astra.
| tier | input / call | output / call | cost / call | 10,000 calls |
|---|---|---|---|---|
| Standard | $0.0400 | $0.1000 | $0.1400 | $1,400 |
| Ultrafast | $0.2400 | $0.6000 | $0.8400 | $8,400 |
The premium is $7,000 for the run. What that buys is the ceiling: up to 6× the speed. If the Standard run takes six hours of wall-clock time, Ultrafast might finish it in about one, so the premium buys roughly five saved hours. The real question is whether those five hours are worth $7,000 to you.
Sometimes they are. If the run blocks a release, feeds a live event, or sits inside a customer-facing promise, five hours can be worth far more than $7,000. Often they are not, and the same job would have been better as a Batch run.
When the trade pays
Ultrafast is rational in a narrow band of situations, and they all share one feature: the deadline is real and the compute is a small part of the value.
- A launch or event where the generation must be ready at a fixed time.
- A run that gates a release, where an extra day of waiting has a cost.
- Interactive bulk generation where a person is watching a long job progress.
- A one-off where the value of the output dwarfs the API bill.
A useful rule of thumb: if the API bill for the run is a rounding error against the value of finishing on time, 6× is cheap. If the run's cost is a line item anyone would notice, 6× is a decision that needs a deadline to justify it. It is also a poor fit for anything you would run again tomorrow, because a repeat run pays the full premium a second time.
When it is pure waste
Most of the time, Ultrafast is the wrong answer, and the alternatives are dramatically cheaper.
- If nobody is waiting, Batch is 0.5× — twelve times cheaper than Ultrafast's 6×.
- If the job is retryable and non-critical, Flex is also 0.5×.
- If a human is waiting but the wait is short, Fast at 2× is a third of the Ultrafast price.
- If the workload has a residency requirement, Ultrafast is unavailable at any price.
Ultrafast and Batch are twelve times apart on the same model. Before paying for speed, check whether the job actually needed to finish today.
The comparison that settles most debates is the last one in that list. If the run can happen overnight, Ultrafast is not a speed purchase — it is twelve times the money for a result you were not going to look at until morning anyway.
Ultrafast and residency
Ultrafast does not support regional processing endpoints, so a workload with a data-residency requirement cannot use it. This is not a surcharge you can absorb; the tier simply is not offered on regional endpoints. A residency-bound product that needs more throughput has to get it from Batch, from a smaller model, or from more parallelism rather than from this tier.
Ultrafast and caching together
Caching still helps on Ultrafast, and it helps proportionally. A cache read is $6.00 per million against $60.00 for uncached input, still a tenth of the rate. On a workload with a large stable prefix, caching removes most of the input cost before the 6× is applied.
Take the same run with a 3,000-token stable prefix inside each 4,000-token input. Sent uncached across 10,000 requests, that is 30 million prefix tokens at $60.00 per million, or $1,800. Cached, the prefix is written once at $75.00 per million and read back at $6.00 per million thereafter, so the same prefix costs a small fraction of $1,800. The speed premium is real, but it does not have to land on tokens you already paid to cache.
Deciding with a number
The Ultrafast decision has a clean form: what is an hour of wall-clock time worth on this job? Price the run both ways, then divide the difference by the hours saved.
A run that takes six hours at Standard and one hour at Ultrafast, costing $7,000 more, is paying about $1,400 per hour saved. If finishing five hours sooner is worth more than $1,400 an hour — because it unblocks a release, meets a fixed deadline, or satisfies a customer promise — the tier is rational. If the work was going to sit in a queue until morning regardless, the same $7,000 buys nothing at all.
Alternatives to paying 6×
Before paying six times the rate, check whether the wall-clock problem can be solved another way.
- Run the job on Batch overnight. Twelve times cheaper, and often fast enough once the deadline stops being same-day.
- Split the work across more parallel requests. Concurrency at Standard rates can beat one faster stream.
- Use a smaller model for the bulk and reserve Astra for the hard cases.
- Shorten the output. Generation is the slow part on long answers, and a tighter maxOutput shortens the run.
Ultrafast is the right tool only when none of those is available and the deadline cannot move.
The short version
- Ultrafast is 6× the Standard rate, applied to input, cache read, cache write and output together.
- Only GPT-6 Astra publishes an Ultrafast price; Sol Ultrafast is announced but not open.
- A 10,000-request bulk run costs $1,400 at Standard and $8,400 at Ultrafast.
- Batch is twelve times cheaper, so Ultrafast is only rational when a real deadline is binding.
- Price it as cost per hour saved, and compare that against the value of finishing early.
- It does not support regional processing, so residency rules it out entirely.
Frequently asked
Which models support Ultrafast?
Only GPT-6 Astra has published Ultrafast pricing. GPT-6.1 Sol Ultrafast has been announced but is not yet open, and no other model offers it.
What does the 6x multiplier apply to?
All four figures: input, cache read, cache write and output. On GPT-6 Astra that moves input from $10.00 to $60.00 per million tokens and output from $50.00 to $300.00.
Is Ultrafast ever cheaper than Batch?
No. Batch is 0.5x the Standard rate and Ultrafast is 6x, a factor of twelve apart. Ultrafast is a latency purchase, never a cost saving.
Does Ultrafast support regional processing?
No. Ultrafast does not support regional processing endpoints, so a residency requirement rules it out regardless of budget.