Skip to main content

Claude Sonnet 5.5 cache price cut: when caching now pays off

Anthropic has halved the price of prompt cache reads on Claude Sonnet 5.5, with cache writes and all other prices unchanged. What changed, why the break-even point stays the same, and how teams running agents on Sonnet should restructure prompts to collect the saving.

Share -
A bearded engineer with his hair tied back studies a monitor showing a node diagram of blue and green data pipeline blocks in a dim office

Anthropic has halved the price of prompt cache reads on Claude Sonnet 5.5, while cache writes and every other price stay the same. Caching still pays off after the same number of reads as before, but each hit after that now costs half as much, so agents that resend a large, stable prefix gain the most.

What was released

A price cut for prompt cache reads on Claude Sonnet 5.5 was announced by Anthropic on 7 October 2026 in the Claude API release notes. Nothing needs to be switched on: requests that already use prompt caching on Sonnet 5.5 are billed at the new rate. It landed on the same day as the Claude Haiku 5.5 launch, which also moved the price picture for teams choosing between Claude models.

What actually changed

Why the break-even point does not move

A cache write costs more than a normal input token, and every later read saves the difference between the input price and the read price. Using Anthropic's published multipliers, the five-minute write premium is recovered by the first read, and the one-hour premium needs two reads. That was true at the old read price and it is still true now, so the cut does not make caching worthwhile for prompts that were not reused before.

What changes is everything after break-even. A coding or support agent that resends the same system prompt, tool definitions and conversation history on every step reads that prefix again on each of those steps. On those requests, the cached part of the input bill now halves, and for long-running agent loops the cached prefix is often most of the input.

What it means for teams building AI products and agents

In CodeDTX's view, these are the practical steps:

  1. Check your hit rate before you count the saving. The response usage fields report cache_creation_input_tokens and cache_read_input_tokens. If both are zero, the prompt was not cached, often because it fell below the minimum. The saving only applies to tokens that actually show up as reads.
  2. Put stable content first. The cache matches the full prefix in order: tools, then system, then messages. Tool definitions and long instructions belong at the front, with anything that changes per request, such as timestamps or user IDs, after the cache breakpoint.
  3. Revisit the one-hour cache. Agents that pause for human approval or wait on slow tools often miss the five-minute window. The one-hour write still needs two reads to pay off, but cheaper reads make a long-lived shared prefix worth more once it is past that point.
  4. Re-run your model comparison. Teams that moved cached, high-volume work from Sonnet to a smaller model on cost grounds should compare again, and teams paying Opus 5.5 rates for heavily cached workloads now have a cheaper Sonnet option to test against their evals. Our guide to controlling what an AI agent costs to run covers how to set that comparison up.
  5. Mind workspace boundaries. On the Claude API each workspace has its own cache, so splitting one product across several workspaces means paying for the same writes more than once.

When the cut does not help

It changes nothing for prompts below the cacheable minimum, for one-off requests with no reuse, or for workloads where output tokens dominate the bill. It does not reduce latency beyond what caching already gave; for that side, see how to make an AI agent respond faster. The release note refers to Claude API pricing, so teams calling Sonnet 5.5 through a cloud provider should confirm the cache-read price on that provider's own pricing page before updating cost budgets.

Frequently asked questions

How much did Anthropic cut the Claude Sonnet 5.5 cache read price?

Anthropic halved it. Cache reads on Claude Sonnet 5.5 now cost half of what they did, moving from one tenth of the base input price to one twentieth, according to the Claude API release notes of 7 October 2026. Cache writes, normal input tokens and output tokens kept their existing prices, so only the cached part of each request gets cheaper.

Do I need to change my code to get the lower cache price?

No. Requests to Claude Sonnet 5.5 that already use prompt caching, either automatic caching or explicit cache breakpoints, are billed at the new read price without any change. It is still worth checking the usage fields in responses to confirm that your prompts are actually being read from the cache, because prompts below the minimum length are processed without caching and return no error.

Does the price cut change when prompt caching pays off?

Not the break-even point. With Anthropic's multipliers, a five-minute cache write is paid back by the first read and a one-hour write by the second, just as before. The difference comes after that: each further read of the same prefix now costs half as much, which matters most for agents that resend long instructions, tool definitions and history on every step.

Is Claude Sonnet 5.5 now cheaper than Opus 5.5 for cached prompts?

Yes. Both models now charge one twentieth of their base input price for a cache read, and because Sonnet 5.5 has the lower base price, its cached tokens cost half as much as Opus 5.5 cached tokens. Before this change they cost the same. Whether Sonnet can do the job is a quality question, so test it against your own evals before switching.

Share this post

Contact us to build the right product

Talk to our engineers about your application, the systems it connects to, and what you want to build next.

Get in touch
Two people discussing work with a laptop