Anthropic has halved the price of prompt cache reads on Claude Sonnet 5.5, while cache writes and every other price stay the same. Caching still pays off after the same number of reads as before, but each hit after that now costs half as much, so agents that resend a large, stable prefix gain the most.
What was released
A price cut for prompt cache reads on Claude Sonnet 5.5 was announced by Anthropic on 7 October 2026 in the Claude API release notes. Nothing needs to be switched on: requests that already use prompt caching on Sonnet 5.5 are billed at the new rate. It landed on the same day as the Claude Haiku 5.5 launch, which also moved the price picture for teams choosing between Claude models.
What actually changed
- Cache reads on Sonnet 5.5 were cut to half their previous price, which Anthropic describes as one twentieth of the base input price instead of one tenth.
- Cache writes and all other prices are unchanged. The prompt caching price table still prices five-minute cache writes above normal input and one-hour writes above that, for every model.
- Sonnet now has the same cache-read multiplier as Opus 5.5. Before the cut, a cached token cost the same on Sonnet 5.5 as on Opus 5.5; it now costs half as much on Sonnet, because Sonnet has the lower base input price. Other models keep the standard multiplier, apart from Fable 5.1 and Mythos 5.1, which are lower still.
- The rules around caching are the same. The minimum cacheable prompt on Sonnet 5.5 is 512 tokens, hits need an identical prefix, and on the Claude API caches are isolated per workspace.
Why the break-even point does not move
A cache write costs more than a normal input token, and every later read saves the difference between the input price and the read price. Using Anthropic's published multipliers, the five-minute write premium is recovered by the first read, and the one-hour premium needs two reads. That was true at the old read price and it is still true now, so the cut does not make caching worthwhile for prompts that were not reused before.
What changes is everything after break-even. A coding or support agent that resends the same system prompt, tool definitions and conversation history on every step reads that prefix again on each of those steps. On those requests, the cached part of the input bill now halves, and for long-running agent loops the cached prefix is often most of the input.
What it means for teams building AI products and agents
In CodeDTX's view, these are the practical steps:
- Check your hit rate before you count the saving. The response
usagefields reportcache_creation_input_tokensandcache_read_input_tokens. If both are zero, the prompt was not cached, often because it fell below the minimum. The saving only applies to tokens that actually show up as reads. - Put stable content first. The cache matches the full prefix in order: tools, then system, then messages. Tool definitions and long instructions belong at the front, with anything that changes per request, such as timestamps or user IDs, after the cache breakpoint.
- Revisit the one-hour cache. Agents that pause for human approval or wait on slow tools often miss the five-minute window. The one-hour write still needs two reads to pay off, but cheaper reads make a long-lived shared prefix worth more once it is past that point.
- Re-run your model comparison. Teams that moved cached, high-volume work from Sonnet to a smaller model on cost grounds should compare again, and teams paying Opus 5.5 rates for heavily cached workloads now have a cheaper Sonnet option to test against their evals. Our guide to controlling what an AI agent costs to run covers how to set that comparison up.
- Mind workspace boundaries. On the Claude API each workspace has its own cache, so splitting one product across several workspaces means paying for the same writes more than once.
When the cut does not help
It changes nothing for prompts below the cacheable minimum, for one-off requests with no reuse, or for workloads where output tokens dominate the bill. It does not reduce latency beyond what caching already gave; for that side, see how to make an AI agent respond faster. The release note refers to Claude API pricing, so teams calling Sonnet 5.5 through a cloud provider should confirm the cache-read price on that provider's own pricing page before updating cost budgets.
Frequently asked questions
How much did Anthropic cut the Claude Sonnet 5.5 cache read price?
Anthropic halved it. Cache reads on Claude Sonnet 5.5 now cost half of what they did, moving from one tenth of the base input price to one twentieth, according to the Claude API release notes of 7 October 2026. Cache writes, normal input tokens and output tokens kept their existing prices, so only the cached part of each request gets cheaper.
Do I need to change my code to get the lower cache price?
No. Requests to Claude Sonnet 5.5 that already use prompt caching, either automatic caching or explicit cache breakpoints, are billed at the new read price without any change. It is still worth checking the usage fields in responses to confirm that your prompts are actually being read from the cache, because prompts below the minimum length are processed without caching and return no error.
Does the price cut change when prompt caching pays off?
Not the break-even point. With Anthropic's multipliers, a five-minute cache write is paid back by the first read and a one-hour write by the second, just as before. The difference comes after that: each further read of the same prefix now costs half as much, which matters most for agents that resend long instructions, tool definitions and history on every step.
Is Claude Sonnet 5.5 now cheaper than Opus 5.5 for cached prompts?
Yes. Both models now charge one twentieth of their base input price for a cache read, and because Sonnet 5.5 has the lower base price, its cached tokens cost half as much as Opus 5.5 cached tokens. Before this change they cost the same. Whether Sonnet can do the job is a quality question, so test it against your own evals before switching.



