OpenAI Fast mode: what 2x the rate actually buys you
Fast is a 2x multiplier for up to 2.5x the speed, bought per request rather than per account. The clean result is that buying Fast for a share of traffic adds roughly that share to the bill.
Fast mode costs twice the Standard rate and buys up to 2.5× the speed. It is the right purchase for a small slice of traffic and the wrong purchase as a default, and the reason is that it is priced per request. You can put 5% of your traffic on Fast, pay for that 5%, and leave the other 95% untouched.
That per-request pricing is the part teams miss. It is easy to hear "2× the rate" and treat Fast as an account-wide switch, when it is really a label you attach to the handful of requests where someone is genuinely watching.
What Fast is
Fast is one of the five service tiers, and it is a straight multiplier: 2× the Standard rate for the same model and the same effort level. The published speed ceiling is up to 2.5×, which is a ceiling rather than a guarantee — latency is a distribution, and Fast moves the whole distribution rather than every individual request by the same factor.
It was called priority until 2026-07-30. The rename is cosmetic; the 2× multiplier and the behaviour are unchanged, so older notes that say "priority" still apply.
It is bought per request
Standard, Batch and Flex are described as properties of the work. Fast is a property of the request. You set it on the calls that need it and nowhere else.
That distinction is the whole reason Fast is affordable. If you have a million requests a month and only 5% of them are on a latency-critical path, you pay 2× on 50,000 requests and 1× on the other 950,000. The premium never touches the bulk of your traffic.
Setting Fast account-wide throws that away. It doubles your entire bill, including the background work that never cared about latency.
The 5% example
Take one million requests a month, an average shape of 2,000 input and 500 output tokens, on GPT-6.1 Sol at Standard rates. The per-call cost is $0.0040 of input plus $0.0050 of output, or $0.0090.
| calls | cost / call | monthly | |
|---|---|---|---|
| Standard | 950,000 | $0.0090 | $8,550 |
| Fast | 50,000 | $0.0180 | $900 |
| Total | 1,000,000 | — | $9,450 |
The bill moves from $9,000 to $9,450, an increase of $450, or exactly 5%. That is not a coincidence. Doubling the rate on a share x of traffic adds x to the total, as long as the requests on Fast have the same token shape as the rest. Buying Fast for 1% of traffic adds about 1%; buying it for 20% adds about 20%.
That single relationship is enough to budget the decision. Before you enable Fast on a route, estimate what fraction of requests will take it. That fraction is your percentage increase.
What 2.5× actually means
The published ceiling is up to 2.5× the speed, and it applies to the request's own latency rather than to throughput. Three consequences follow.
- Fast does not make a slow request fast. It moves a request up the queue; the model still has to generate the tokens.
- The benefit is largest when queueing is the bottleneck, and smallest when generation time dominates.
- Because it is a distribution shift, measure your own p50 and p95 rather than trusting the headline multiple.
For a short request where queueing is most of the wait, Fast can be transformative. For a long generation, the ceiling is doing less work than the marketing suggests.
Fast and regional residency
Fast does not support regional processing endpoints. This is a hard constraint, not a surcharge you can pay around. If a request must be processed in a specific region, it cannot be served on Fast at any price, and the +10% regional surcharge applies only to the tiers that do offer residency — Standard, Batch and Flex.
The practical result is that a residency-bound product with a latency problem cannot solve it with a tier. The latency has to come out of the prompt, the model choice, or the architecture.
When to buy it
Fast earns its 2× in a narrow set of situations, all of them sharing one property: a human is waiting and the wait is visible.
- A conversational surface where the perceived delay changes how the product feels.
- A request a user has explicitly triggered and is watching.
- A step in a chain where each request's latency compounds into a total a user experiences.
- A retry after a timeout, where the alternative is failing the user.
It does not earn its 2× for background jobs, scheduled work, anything already asynchronous, or any request whose result is consumed by a process rather than a person. Those belong on Standard, or on Batch and Flex if nobody is watching at all.
Fast is a per-request upgrade, not a plan. If you cannot name the specific route that needs it, you do not need it yet.
How to tell whether Fast is working
Fast is easy to enable and hard to evaluate, because latency is noisy. Before turning it on for a route, record the p50 and p95 latency for that route over a normal week. Then enable it and measure again. If the p95 barely moves, queueing was not the bottleneck and you are paying 2× for very little.
Two failure patterns are common. The first is a route where generation time dominates, so a speed ceiling has little to work with. The second is a route where your own stack — a database call, a retrieval step, a downstream service — accounts for most of the wait, in which case a faster model does not make the user's wait shorter.
Cheaper ways to cut latency first
Doubling the rate is the most expensive latency lever, so exhaust the cheap ones before pulling it.
- Trim the input. Fewer tokens in means less to process, and it costs nothing.
- Cap maxOutput. A model that stops at 300 tokens returns sooner than one that rambles to 900.
- Cache the stable prefix. A cache read is cheaper and faster than reprocessing the same tokens.
- Lower effort on tasks that do not need reasoning. Less thinking is less generation.
- Move the call off the critical path. If the user does not need the result immediately, Standard or Batch is the honest answer.
| lever | cost | effect on the wait |
|---|---|---|
| Trim the input | free | less to process |
| Cap maxOutput | free | shorter generation |
| Cache the prefix | below input rate | skips reprocessing |
| Lower effort | lower output cost | less generation |
| Fast tier | 2× | higher queue priority |
| Ultrafast tier | 6× | higher priority, Astra only |
Only when those are done does the 2× multiplier buy something the architecture cannot.
Fast and the rest of the latency budget
Fast shifts one component of the user's wait: the model call. It does nothing for the network, the retrieval step, the database, or the client. A product whose model call is a fifth of the total latency will see a small end-to-end improvement even if Fast delivers exactly what it promises on that slice.
That is the argument for measuring end-to-end rather than model-side. The tier is priced per request, and it is worth 2× only when the model call is a large enough share of the wait that shortening it changes what the user feels. When it is not, the money is better spent removing the other four-fifths.
The short version
- Fast is 2× the Standard rate for up to 2.5× the speed, formerly called priority.
- It is priced per request, so the premium applies only to the calls you mark.
- Doubling the rate on a share x of traffic adds about x to the bill; 5% of traffic costs 5% more.
- Fast does not support regional processing, so residency rules it out entirely.
- Measure p50 and p95 before and after; trim input and cap output before you pay 2×.
- Buy it only for requests a human is watching, and nowhere else.
Frequently asked
What is Fast mode called now?
Fast. It was named priority until 2026-07-30, and the multiplier did not change when the name did.
Does Fast mode make the model smarter?
No. It buys speed, not quality. The model and the effort level are unchanged; only the scheduling priority and the price are different.
Can I enable Fast for my whole account?
You can, but you should not. Fast is priced per request, so the sensible pattern is to mark the small set of latency-critical requests and leave everything else on Standard.
Does Fast support regional processing?
No. Fast does not support regional processing endpoints, so a request with a residency requirement cannot use it.