Caching

OpenAI prompt caching: what cached input actually costs

Cache reads cost 5% to 10% of the input rate and cache writes cost 1.25x, so a reused prefix pays for itself on the second call. The ratio survives the long-context cliff.

Prompt caching has two prices, and they are not the same number. Cached input is billed far below the normal input rate; writing to the cache costs slightly above it. Get those two figures straight and caching stops being a vague optimisation and becomes a break-even question with an unusually short answer: the second call already pays for the write.

That short break-even is why caching is worth doing almost anywhere a prefix repeats. It is not a trick for special workloads. It is the default shape of any application that sends the same instructions, tools or context more than once.

The two cache prices

Every model on the price list has a cache read rate and a cache write rate alongside its input and output rates. The read rate is a small fraction of input; the write rate is a fixed 1.25× on every model.

modelinputcache readread as % of inputcache writewrite multiple
GPT-6.1 Sol$2.00$0.105%$2.501.25×
GPT-6 Astra$10.00$1.0010%$12.501.25×
GPT-6 Luna$0.10$0.0110%$0.1251.25×
GPT-5.6 Sol$4.00$0.4010%$5.001.25×
GPT-5.6 Terra$2.00$0.2010%$2.501.25×

Rates are USD per million tokens. GPT-6.1 Sol is the outlier at 5%, which makes caching on that model slightly more attractive than on the rest. Everywhere else the read is a tenth of the input price, which is still a very large discount.

Break-even on a cache write

Take a 10,000-token prefix sent on GPT-6.1 Sol, and compare sending it uncached every time with writing it once and reading it back.

Without caching, each call pays 10,000 × $2.00 per million, or $0.020. With caching, the first call pays the write rate, 10,000 × $2.50 per million, or $0.025, and every later call pays the read rate, 10,000 × $0.10 per million, or $0.001.

callsuncached prefixcached prefixsaving
1$0.020$0.025−25%
2$0.040$0.02635%
5$0.100$0.02971%
10$0.200$0.03483%
100$2.000$0.12494%

The write premium is $0.005, and each cached read saves $0.019 against the uncached rate. The break-even is therefore $0.005 ÷ $0.019, or about 1.26 reads. From the second call onwards, caching is ahead, and by the tenth call it has removed 83% of the prefix cost.

The asymmetry is worth holding on to: the write is a one-off and the reads are recurring, so the more a prefix repeats, the closer the average cost gets to the read rate. At a hundred calls the prefix costs 6% of what it would uncached, which is close to the floor the read rate sets.

Long context: the ratio survives

Crossing the 272,000-token threshold re-prices the whole request: input ×2, cache read ×2, cache write ×2 and output ×1.5. Because the input rate and the cache read rate are both doubled, the ratio between them does not change. On GPT-6.1 Sol the cached read is still 5% of the input rate after the cliff.

That matters because it is tempting to assume the cliff wipes out the caching advantage. It does not. Take a 280,000-token request of which 250,000 tokens are a stable cached prefix. After the long-context multiplier, the input rate is $4.00 per million and the cache read rate is $0.20 per million. The 250,000-token prefix costs $0.05 cached against $1.00 uncached — still a 95% saving on that portion.

The cliff and caching are independent. Caching will not pull you back under the threshold, because it does not change the token count, but it stays just as valuable once you are over it.

Keep the prefix stable

Caching works on exact prefixes. Any change early in the prompt invalidates everything after it, so the layout of the request matters as much as the content.

A prefix that changes once a day is effectively static and caches well. A prefix that changes once a request is not a prefix at all. A useful test is to ask how often the early part of the prompt changes; if the answer is "never during a session", it should be cached, and if it is "every call", the caching work belongs somewhere else in the system.

Caching in an agent loop

Agent loops are where caching earns the most and where it is most often missed. History grows every turn, so an uncached loop pays for the entire conversation again on every step — and it approaches the 272K cliff much sooner than it needs to.

With a stable cached prefix, each turn only pays the read rate for the history plus the full rate for the new tokens. The saving compounds across the loop, and the effective context budget stretches further because the cached portion is cheap. The two pieces of advice reinforce each other: keep the prefix stable, and the loop stays cheaper and under the cliff for longer.

One caution: changing effort between requests can change the rendered prompt and drop the cache. If you vary effort per turn and also rely on caching, check that the change does not sit inside the cached prefix.

Where caching saves the most

The saving scales with how much of the prompt is stable and how often it repeats. Four shapes get the most.

Where it does not help

Caching is not free money, and there are prompts it does nothing for.

Caching stacks with the tiers

Cache rates are multipliers on the model's rate, so the tier multiplies them too. In Batch or Flex a cache read is $0.05 per million on Sol rather than $0.10, and a cache write is $1.25 rather than $2.50. On the fast tiers the same figures double instead.

That means the cheapest configuration for a large repeated prefix is usually a half-price tier with a cached prefix, not either lever on its own. The two compose, and neither cancels the other.

The short version

Frequently asked

How much does a cached input token cost?

A fraction of the normal input rate: 0.10 against 2.00 per million on GPT-6.1 Sol, which is 5%, and 10% on GPT-6 Astra, GPT-6 Luna and both GPT-5.6 models.

What does writing to the cache cost?

1.25x the input rate on every model. On GPT-6.1 Sol a cache write is $2.50 per million tokens against $2.00 for a normal input token.

How many times must I reuse a prefix before caching pays?

Twice. The cache write premium is 0.25x and each cached read saves 0.90x against the input rate, so the break-even is about 1.26 reads — the second call is already ahead.

Does caching still help above the 272K threshold?

Yes. The long-context multiplier doubles both the input rate and the cache read rate, so the ratio between them is unchanged and the relative saving on a cached prefix is the same.