Long context

The 272K long-context cliff, and how to stay on the right side of it

The long-context threshold is a cliff, not a slope: crossing 272,000 input tokens roughly doubles the cost of the entire request. Two worked examples and four ways to stay underneath.

The 272,000-token threshold is not a taper. When input crosses it, the entire request is re-priced: input, cache read and cache write at 2×, and output at 1.5×. The multiplier does not apply to the tokens above the line. It applies to all of them. One token over the threshold can roughly double the cost of the whole request.

That is why this site exists, and why the number is worth understanding before you tune anything else. A long-context request that sits just under the line is not slightly cheaper than one just over it. It is about half the price.

It is a cliff, not a slope

The rule is simple to state and easy to misread. The threshold is 272,000 input tokens. A request at or below it pays the standard rates. A request above it pays standard rates multiplied by the long-context factors for the whole request:

figuremultiplier above 272K
input×2
cache read×2
cache write×2
output×1.5

The intuition most people bring is that only the excess is charged more, the way a progressive tax works. That is not the case here. The moment the input count crosses the line, every input token in the request is billed at double, including the first one. The cliff is a discontinuity, and the only way to see it is to price the same request on both sides.

GPT-6.1 Sol at the boundary

Take a request with 272,000 input tokens and 2,000 output tokens on GPT-6.1 Sol. At the base rates, input costs 272,000 × $2.00 per million, or $0.544, and output costs 2,000 × $10.00 per million, or $0.020. The total is $0.564.

Now add a single token. At 272,001 input tokens the request is over the threshold, so input is billed at $4.00 per million — 272,001 × $4.00 per million, or $1.088 — and output at $15.00 per million, or $0.030. The total is $1.118.

input costoutput costtotal
272,000 input tokens$0.544$0.020$0.564
272,001 input tokens$1.088$0.030$1.118

One extra token takes the bill from $0.564 to $1.118, a factor of 1.98. The token itself is worth a fraction of a cent; the re-pricing is worth fifty-four cents.

GPT-6 Astra at the boundary

The same shape appears on Astra, scaled by its higher rates. At 272,000 input tokens and 2,000 output tokens, input costs $2.720 and output $0.100, for $2.820. One token over, input is billed at $20.00 per million and output at $75.00 per million, giving $5.440 and $0.150, for $5.590.

input costoutput costtotal
272,000 input tokens$2.720$0.100$2.820
272,001 input tokens$5.440$0.150$5.590

Again the ratio is 1.98. The absolute jump is larger because Astra's rates are larger, but the mechanism is identical: the whole request is re-priced, and input dominates the total, so the bill nearly doubles.

Four ways back under 272K

If a request is just over the line, the cheapest fix is almost always to remove input rather than to accept the multiplier. Four approaches, in rough order of how much they usually save:

Caching and the cliff are independent. A cached prefix still counts towards the input total, so caching alone will not rescue a request that is over the threshold. Measure the input count directly.

The order matters. Trimming retrieval usually removes more tokens than summarising history and is easier to justify, because a smaller retrieval set is often a better answer rather than a worse one. Chunking is the fallback when the content genuinely cannot be reduced.

When crossing is correct

The cliff is a price, not a prohibition, and there are requests where paying it is the right call.

What is never correct is crossing by accident — twenty retrieved chunks where three would do, or an agent loop that carries its entire history because nobody summarised it. The multiplier is defensible when it is chosen and expensive when it is discovered.

Why output is multiplied by less

The long-context factors are not uniform: input, cache read and cache write are doubled, but output is multiplied by 1.5. That asymmetry changes which requests cross cheaply.

A request that is almost all input — a retrieval lookup over a large document with a short answer — nearly doubles when it crosses. A request that is mostly output — a long generation conditioned on a moderate prompt — crosses more gently, because the expensive column is only multiplied by 1.5. If you have to cross, cross with an output-heavy request.

Measure before you send

The threshold is on the input count, so the only reliable way to stay under it is to count before you send. Token counts vary with formatting, so an estimate from character length is not good enough near the boundary.

The cliff and the half-price tiers

The long-context multiplier composes with the tier multiplier, and the order still matters. A request over the cliff on Batch is 0.5× of a doubled number; the same request trimmed below the cliff and then batched is 0.5× of the base number. Trimming first and batching second is the cheaper sequence, as the worked example in the Batch guide shows.

Caching does not change the input count, so it cannot move a request below the line. What it can do is make the stable part of a long prompt cheap enough that the doubled rate hurts less — useful, but not a substitute for trimming.

The short version

Frequently asked

What happens when input exceeds 272,000 tokens?

The whole request is re-priced, not just the excess. Input, cache read and cache write are multiplied by 2 and output by 1.5, which roughly doubles the bill for one extra token.

Is the multiplier applied only to the tokens above 272K?

No. This is the common misunderstanding. The multiplier applies to the entire request, which is why the threshold behaves like a cliff rather than a gradual increase.

Does the threshold apply to cache reads as well?

Yes. Cache read and cache write are both multiplied by 2 above the threshold, so cached prefixes are re-priced too, though their relative advantage is unchanged.

Is it ever correct to cross the threshold?

Yes, when the task genuinely needs the whole document and splitting it would lose meaning. Output-heavy requests also cross more cheaply, because output is multiplied by 1.5 rather than 2.