Methodology — where every number on this site comes from
Every rate, multiplier and assumption used on this site, separated into what OpenAI publishes, what we estimate, and what we could not confirm.
This page exists so that no number on the site has to be taken on trust. It separates what OpenAI publishes from what we assume, gives the formula the calculator actually runs, and lists the questions we have not been able to settle.
The three kinds of number on this site
Published rates and multipliers. Token prices, cache prices, the long-context multipliers, the service-tier multipliers and the regional surcharge are read from OpenAI's API pricing documentation. These are facts with a date attached.
Planning assumptions. Anything about how much work a reasoning effort level causes is an assumption, because OpenAI publishes no fixed token multiplier per level. Every effort multiplier in the calculator is editable for that reason.
Unconfirmed figures. A small number of interactions are not spelled out in the documentation. They are listed under "Open questions" below and handled conservatively in the calculator.
Rate sources
| Source | What it gives us | Verified |
|---|---|---|
developers.openai.com/api/docs/pricing | Base rates, cache rates, long-context bands, tier multipliers | 2026-10-01 |
| Published Ultrafast breakdown for GPT-6 Astra | The 6x multiplier on all four figures | 2026-10-01 |
| A third-party OpenAI pricing table | Cross-check on long-context multipliers | 2026-10-01 |
Rates are stored in a single data file that every page reads from. That file carries the verification date, and the footer of every page renders it, so a stale figure is visible rather than silent.
The formula
A request is priced in three steps. First, the base rates for the model. Then a long-context multiplier applied to the whole request if input crosses the threshold. Then the service-tier multiplier, and the regional surcharge if applicable:
step 1 base = input x input_rate + cache_read x cache_read_rate
+ cache_write x cache_write_rate + output x output_rate
step 2 if input_tokens > 272,000:
every term above is multiplied by its long-context factor
(input x2, cache read x2, cache write x2, output x1.5)
step 3 multiply the result by the service-tier multiplier
(batch 0.5, flex 0.5, standard 1, fast 2, ultrafast 6)
and by 1.10 if the request runs on a regional endpoint
The step-2 rule is the one people get wrong. The multiplier applies to the entire request, not only to the tokens above the threshold. A request at 273,000 input tokens is priced as though all 273,000 crossed the line.
Service-tier multipliers
| Tier | Multiplier | Speed ceiling | Regional residency |
|---|---|---|---|
| Batch | 0.5 | up to 24 hours | yes |
| Flex | 0.5 | queued | yes |
| Standard | 1 | 1x | yes |
| Fast | 2 | up to 2.5x | no |
| Ultrafast | 6 | up to 6x | no |
Fast was previously called priority and was renamed in July 2026. Ultrafast is currently available for GPT-6 Astra only. Speed ceilings are the published maximum, not a guarantee.
Effort multipliers
| Level | Thinking-token multiplier | Status |
|---|---|---|
low | 1 | planning assumption |
medium | 3 | planning assumption |
high | 8 | planning assumption |
xhigh | 20 | planning assumption |
max | 45 | planning assumption |
These are a starting point for routing decisions, not a published specification. If you have real traces, measure your own ratios and replace them. The arithmetic on this site does not change when you do — only the column values move.
Open questions
Three figures are not spelled out clearly enough in the public documentation for us to state them as fact. The calculator takes the conservative reading in each case:
- Whether the Batch and Flex discounts apply to cached input and cache writes, or only to the standard token lines. We apply the tier multiplier to the whole request.
- Whether the higher cache-write rate for longer cache lifetimes applies to every current model. We use the published figure per model and flag it here rather than assume uniformity.
- Which rate-limit tier Ultrafast requests are served from. We do not model rate limits at all.
If you have documentation that settles any of these, send it and this page will change.
What the calculator does not model
- Retries and failures. A request that fails and is re-sent costs twice. The calculator counts successful requests only.
- Rate limits and queueing. Flex may queue or be unavailable; that risk is a description, not a number.
- Storage or hosting costs for the data you batch.
- Latency. The calculator is about money. Faster tiers also cost more, which is a separate trade-off.
Corrections
If a rate on this site is out of date, or a table disagrees with the calculator, tell us. Corrections are applied to the page itself and the "updated" date changes, so there is one current version rather than an accumulating errata list. See the contact page.
Independence
TierCliff is not affiliated with, endorsed by or sponsored by OpenAI. Product names and prices referenced on this site belong to their respective owners. The site is funded by advertising, and no advertiser has any influence over the recommendations the tool makes.