The 272K long-context cliff, and how to stay on the right side of it
The long-context threshold is a cliff, not a slope: crossing 272,000 input tokens roughly doubles the cost of the entire request. Two worked examples and four ways to stay underneath.
The 272,000-token threshold is not a taper. When input crosses it, the entire request is re-priced: input, cache read and cache write at 2×, and output at 1.5×. The multiplier does not apply to the tokens above the line. It applies to all of them. One token over the threshold can roughly double the cost of the whole request.
That is why this site exists, and why the number is worth understanding before you tune anything else. A long-context request that sits just under the line is not slightly cheaper than one just over it. It is about half the price.
It is a cliff, not a slope
The rule is simple to state and easy to misread. The threshold is 272,000 input tokens. A request at or below it pays the standard rates. A request above it pays standard rates multiplied by the long-context factors for the whole request:
| figure | multiplier above 272K |
|---|---|
| input | ×2 |
| cache read | ×2 |
| cache write | ×2 |
| output | ×1.5 |
The intuition most people bring is that only the excess is charged more, the way a progressive tax works. That is not the case here. The moment the input count crosses the line, every input token in the request is billed at double, including the first one. The cliff is a discontinuity, and the only way to see it is to price the same request on both sides.
GPT-6.1 Sol at the boundary
Take a request with 272,000 input tokens and 2,000 output tokens on GPT-6.1 Sol. At the base rates, input costs 272,000 × $2.00 per million, or $0.544, and output costs 2,000 × $10.00 per million, or $0.020. The total is $0.564.
Now add a single token. At 272,001 input tokens the request is over the threshold, so input is billed at $4.00 per million — 272,001 × $4.00 per million, or $1.088 — and output at $15.00 per million, or $0.030. The total is $1.118.
| input cost | output cost | total | |
|---|---|---|---|
| 272,000 input tokens | $0.544 | $0.020 | $0.564 |
| 272,001 input tokens | $1.088 | $0.030 | $1.118 |
One extra token takes the bill from $0.564 to $1.118, a factor of 1.98. The token itself is worth a fraction of a cent; the re-pricing is worth fifty-four cents.
GPT-6 Astra at the boundary
The same shape appears on Astra, scaled by its higher rates. At 272,000 input tokens and 2,000 output tokens, input costs $2.720 and output $0.100, for $2.820. One token over, input is billed at $20.00 per million and output at $75.00 per million, giving $5.440 and $0.150, for $5.590.
| input cost | output cost | total | |
|---|---|---|---|
| 272,000 input tokens | $2.720 | $0.100 | $2.820 |
| 272,001 input tokens | $5.440 | $0.150 | $5.590 |
Again the ratio is 1.98. The absolute jump is larger because Astra's rates are larger, but the mechanism is identical: the whole request is re-priced, and input dominates the total, so the bill nearly doubles.
Four ways back under 272K
If a request is just over the line, the cheapest fix is almost always to remove input rather than to accept the multiplier. Four approaches, in rough order of how much they usually save:
- Trim retrieval. Return fewer, better chunks. A retrieval step that returns twenty documents when three would answer the question is the most common cause of an accidental cliff crossing, and cutting it is free.
- Summarise history. In a long conversation or agent loop, collapse older turns into a running summary. The recent turns stay verbatim; the distant ones become a paragraph.
- Chunk and join. Split the work into two requests that each stay under the threshold, then combine the answers in a second, short call. The join costs a fraction of the doubled request.
- Restructure the prompt and cache the stable prefix. Move instructions and tool definitions into a stable prefix that is cached, and keep the volatile material short. Caching does not change the token count, so it cannot by itself move you under the line, but it makes the stable part cheap and frees you to trim the volatile part harder.
Caching and the cliff are independent. A cached prefix still counts towards the input total, so caching alone will not rescue a request that is over the threshold. Measure the input count directly.
The order matters. Trimming retrieval usually removes more tokens than summarising history and is easier to justify, because a smaller retrieval set is often a better answer rather than a worse one. Chunking is the fallback when the content genuinely cannot be reduced.
When crossing is correct
The cliff is a price, not a prohibition, and there are requests where paying it is the right call.
- The task genuinely needs the whole document in view, and splitting it would break the reasoning. Legal review, full-codebase refactors and whole-report analysis are the usual examples.
- The alternative is a worse answer. If missing a clause is unacceptable, the doubled price is cheap.
- The request is output-heavy. Output is multiplied by 1.5 rather than 2, so a request dominated by generation crosses more cheaply than one dominated by input.
- You have already trimmed and the remainder is the minimum. At that point the multiplier is the honest price of the task.
What is never correct is crossing by accident — twenty retrieved chunks where three would do, or an agent loop that carries its entire history because nobody summarised it. The multiplier is defensible when it is chosen and expensive when it is discovered.
Why output is multiplied by less
The long-context factors are not uniform: input, cache read and cache write are doubled, but output is multiplied by 1.5. That asymmetry changes which requests cross cheaply.
A request that is almost all input — a retrieval lookup over a large document with a short answer — nearly doubles when it crosses. A request that is mostly output — a long generation conditioned on a moderate prompt — crosses more gently, because the expensive column is only multiplied by 1.5. If you have to cross, cross with an output-heavy request.
Measure before you send
The threshold is on the input count, so the only reliable way to stay under it is to count before you send. Token counts vary with formatting, so an estimate from character length is not good enough near the boundary.
- Count the assembled prompt, not the pieces. Retrieval, history, instructions and tool definitions all add up.
- Leave headroom. Aim for a comfortable margin below 272,000 rather than landing at 271,990.
- Log the input count per request, so a prompt change that pushes traffic over the line shows up in the data rather than the invoice.
The cliff and the half-price tiers
The long-context multiplier composes with the tier multiplier, and the order still matters. A request over the cliff on Batch is 0.5× of a doubled number; the same request trimmed below the cliff and then batched is 0.5× of the base number. Trimming first and batching second is the cheaper sequence, as the worked example in the Batch guide shows.
Caching does not change the input count, so it cannot move a request below the line. What it can do is make the stable part of a long prompt cheap enough that the doubled rate hurts less — useful, but not a substitute for trimming.
The short version
- Crossing 272,000 input tokens re-prices the whole request, not the excess.
- Input, cache read and cache write are multiplied by 2; output by 1.5.
- On both Sol and Astra, one token over the line takes the bill to 1.98×.
- Trim retrieval first, then summarise history, then chunk and join.
- Caching does not reduce the token count, so it cannot pull you back under the line.
- Cross deliberately for whole-document tasks, never by accident.
Frequently asked
What happens when input exceeds 272,000 tokens?
The whole request is re-priced, not just the excess. Input, cache read and cache write are multiplied by 2 and output by 1.5, which roughly doubles the bill for one extra token.
Is the multiplier applied only to the tokens above 272K?
No. This is the common misunderstanding. The multiplier applies to the entire request, which is why the threshold behaves like a cliff rather than a gradual increase.
Does the threshold apply to cache reads as well?
Yes. Cache read and cache write are both multiplied by 2 above the threshold, so cached prefixes are re-priced too, though their relative advantage is unchanged.
Is it ever correct to cross the threshold?
Yes, when the task genuinely needs the whole document and splitting it would lose meaning. Output-heavy requests also cross more cheaply, because output is multiplied by 1.5 rather than 2.