OpenAI prompt caching: what cached input actually costs
Cache reads cost 5% to 10% of the input rate and cache writes cost 1.25x, so a reused prefix pays for itself on the second call. The ratio survives the long-context cliff.
Prompt caching has two prices, and they are not the same number. Cached input is billed far below the normal input rate; writing to the cache costs slightly above it. Get those two figures straight and caching stops being a vague optimisation and becomes a break-even question with an unusually short answer: the second call already pays for the write.
That short break-even is why caching is worth doing almost anywhere a prefix repeats. It is not a trick for special workloads. It is the default shape of any application that sends the same instructions, tools or context more than once.
The two cache prices
Every model on the price list has a cache read rate and a cache write rate alongside its input and output rates. The read rate is a small fraction of input; the write rate is a fixed 1.25× on every model.
| model | input | cache read | read as % of input | cache write | write multiple |
|---|---|---|---|---|---|
| GPT-6.1 Sol | $2.00 | $0.10 | 5% | $2.50 | 1.25× |
| GPT-6 Astra | $10.00 | $1.00 | 10% | $12.50 | 1.25× |
| GPT-6 Luna | $0.10 | $0.01 | 10% | $0.125 | 1.25× |
| GPT-5.6 Sol | $4.00 | $0.40 | 10% | $5.00 | 1.25× |
| GPT-5.6 Terra | $2.00 | $0.20 | 10% | $2.50 | 1.25× |
Rates are USD per million tokens. GPT-6.1 Sol is the outlier at 5%, which makes caching on that model slightly more attractive than on the rest. Everywhere else the read is a tenth of the input price, which is still a very large discount.
Break-even on a cache write
Take a 10,000-token prefix sent on GPT-6.1 Sol, and compare sending it uncached every time with writing it once and reading it back.
Without caching, each call pays 10,000 × $2.00 per million, or $0.020. With caching, the first call pays the write rate, 10,000 × $2.50 per million, or $0.025, and every later call pays the read rate, 10,000 × $0.10 per million, or $0.001.
| calls | uncached prefix | cached prefix | saving |
|---|---|---|---|
| 1 | $0.020 | $0.025 | −25% |
| 2 | $0.040 | $0.026 | 35% |
| 5 | $0.100 | $0.029 | 71% |
| 10 | $0.200 | $0.034 | 83% |
| 100 | $2.000 | $0.124 | 94% |
The write premium is $0.005, and each cached read saves $0.019 against the uncached rate. The break-even is therefore $0.005 ÷ $0.019, or about 1.26 reads. From the second call onwards, caching is ahead, and by the tenth call it has removed 83% of the prefix cost.
The asymmetry is worth holding on to: the write is a one-off and the reads are recurring, so the more a prefix repeats, the closer the average cost gets to the read rate. At a hundred calls the prefix costs 6% of what it would uncached, which is close to the floor the read rate sets.
Long context: the ratio survives
Crossing the 272,000-token threshold re-prices the whole request: input ×2, cache read ×2, cache write ×2 and output ×1.5. Because the input rate and the cache read rate are both doubled, the ratio between them does not change. On GPT-6.1 Sol the cached read is still 5% of the input rate after the cliff.
That matters because it is tempting to assume the cliff wipes out the caching advantage. It does not. Take a 280,000-token request of which 250,000 tokens are a stable cached prefix. After the long-context multiplier, the input rate is $4.00 per million and the cache read rate is $0.20 per million. The 250,000-token prefix costs $0.05 cached against $1.00 uncached — still a 95% saving on that portion.
The cliff and caching are independent. Caching will not pull you back under the threshold, because it does not change the token count, but it stays just as valuable once you are over it.
Keep the prefix stable
Caching works on exact prefixes. Any change early in the prompt invalidates everything after it, so the layout of the request matters as much as the content.
- Put static material first: system instructions, tool definitions, reference documents.
- Put the variable material last: the user's question, the current turn, request-specific data.
- Keep timestamps, request IDs and session tokens out of the system prompt, or they will change the prefix on every call.
- Treat the prefix as an interface. If it changes often, it was not really static.
A prefix that changes once a day is effectively static and caches well. A prefix that changes once a request is not a prefix at all. A useful test is to ask how often the early part of the prompt changes; if the answer is "never during a session", it should be cached, and if it is "every call", the caching work belongs somewhere else in the system.
Caching in an agent loop
Agent loops are where caching earns the most and where it is most often missed. History grows every turn, so an uncached loop pays for the entire conversation again on every step — and it approaches the 272K cliff much sooner than it needs to.
With a stable cached prefix, each turn only pays the read rate for the history plus the full rate for the new tokens. The saving compounds across the loop, and the effective context budget stretches further because the cached portion is cheap. The two pieces of advice reinforce each other: keep the prefix stable, and the loop stays cheaper and under the cliff for longer.
One caution: changing effort between requests can change the rendered prompt and drop the cache. If you vary effort per turn and also rely on caching, check that the change does not sit inside the cached prefix.
Where caching saves the most
The saving scales with how much of the prompt is stable and how often it repeats. Four shapes get the most.
- A fixed system prompt with tool definitions, sent on every request. On a 6,000-token instruction block, the uncached cost is $0.012 per call on Sol; cached, it is $0.0006 after the first write.
- Few-shot examples that never change. They are the expensive part of many classification prompts and the easiest thing to cache.
- An agent loop carrying a growing history. The prefix stays stable while only the newest turn changes, so each step pays the read rate for everything before it.
- A RAG prompt whose retrieved documents are stable for a session. Cache the instructions and the fixed corpus, and let only the question vary.
Where it does not help
Caching is not free money, and there are prompts it does nothing for.
- One-shot requests. If a prefix is sent once, you pay the 1.25× write premium and never read it back.
- A prefix that changes every call. Timestamps, request IDs and per-user data early in the prompt make each request a fresh write.
- A tiny prefix. Below a few hundred tokens the absolute saving is fractions of a cent, and the write premium can exceed it.
- A prompt whose layout changes when effort changes, because that can drop the cache on every turn.
Caching stacks with the tiers
Cache rates are multipliers on the model's rate, so the tier multiplies them too. In Batch or Flex a cache read is $0.05 per million on Sol rather than $0.10, and a cache write is $1.25 rather than $2.50. On the fast tiers the same figures double instead.
That means the cheapest configuration for a large repeated prefix is usually a half-price tier with a cached prefix, not either lever on its own. The two compose, and neither cancels the other.
The short version
- Cached input costs 5% of the input rate on GPT-6.1 Sol and 10% on the other models.
- A cache write costs 1.25× the input rate, a premium of 0.25×.
- Break-even is about 1.26 reads, so the second call is already ahead.
- The long-context cliff doubles both rates, so the caching advantage is unchanged.
- Caching does not reduce token count, so it cannot pull a request back under 272K.
- Keep the static prefix first and the variable part last, or you will invalidate it every call.
Frequently asked
How much does a cached input token cost?
A fraction of the normal input rate: 0.10 against 2.00 per million on GPT-6.1 Sol, which is 5%, and 10% on GPT-6 Astra, GPT-6 Luna and both GPT-5.6 models.
What does writing to the cache cost?
1.25x the input rate on every model. On GPT-6.1 Sol a cache write is $2.50 per million tokens against $2.00 for a normal input token.
How many times must I reuse a prefix before caching pays?
Twice. The cache write premium is 0.25x and each cached read saves 0.90x against the input rate, so the break-even is about 1.26 reads — the second call is already ahead.
Does caching still help above the 272K threshold?
Yes. The long-context multiplier doubles both the input rate and the cache read rate, so the ratio between them is unchanged and the relative saving on a cached prefix is the same.