Skip to main content

Claude Haiku 5.5: lower prices and the API changes to fix first

Anthropic's Claude Haiku 5.5 cuts per-token prices for short prompts and adds a 1M token context, but code written for Haiku 4.5 can break. What changed, what to fix before switching, and when to stay put.

Share -
A person holds a smartphone in front of a glowing blue holographic API emblem surrounded by connected app and user icons on a dark grid

Claude Haiku 5.5 is Anthropic's new small model for high-volume, latency-sensitive work such as classification, routing, extraction and subagents. It is much cheaper per token than Haiku 4.5 on short prompts and adds a 1M token context, but it is not a drop-in swap: several request patterns that worked on Haiku 4.5 now return errors.

What was released

Claude Haiku 5.5 was announced by Anthropic on 7 October 2026 in the Claude API release notes, under the model ID claude-haiku-5-5. It is available on the Claude API, Claude in Amazon Bedrock, Claude Platform on AWS, Claude on Google Cloud and Claude in Microsoft Foundry. The ID is fixed, with no date suffix and no separate alias, and each platform's form is listed in the migration guide.

What actually changed

The What's new in Claude Haiku 5.5 page and the pricing page list these differences from Haiku 4.5:

  • Lower prices for short prompts, a step up for long ones. For prompts up to 100,000 tokens, Haiku 5.5's list prices for input and output are a tenth of Haiku 4.5's. Prompts over 100,000 tokens pay a higher rate, which is still half of Haiku 4.5's list prices. Haiku 5.5 is the only current Claude model priced by prompt length.
  • A bigger context and longer output. The context window grows from 200k to 1M tokens, and maximum output from 64k to 128k tokens.
  • Adaptive thinking, on by default. The model decides when and how much to think, and the effort parameter is the main lever for trading quality against speed and cost.
  • A new tokenizer. The same text counts as more tokens, roughly 1.3 for every one on Haiku 4.5, so usage figures, token budgets and cost estimates all move.
  • Browser use. Haiku 5.5 supports the browser use tool on the Claude API and Google Cloud, which Haiku 4.5 does not.

What breaks if you only change the model ID

The migration guide lists five breaking changes. Each returns a 400 error rather than a quietly different answer, so a test run against your real requests will find them:

  • thinking with budget_tokens is rejected. Use adaptive thinking and set effort instead.
  • temperature, top_p and top_k at non-default values are rejected. Steer with the prompt.
  • An assistant prefill at the end of messages is rejected, even with thinking off. Use structured outputs or tools for format, and move continuations into the user turn.
  • Computer use on the Claude API and Google Cloud needs the computer_toolset_20260801 toolset in place of computer_20250124.
  • Sending thinking blocks back after changing system, tools or earlier messages is rejected, so conversations that carry thinking must stay append-only.

Two quieter changes catch code that never errors. A response can start with a thinking block, so code that reads the first content block as the answer must select by type. And thinking tokens count toward max_tokens, so a small limit tuned for Haiku 4.5 can stop after the thinking and before any text.

Anthropic notes that agents built on Claude Managed Agents need only the model name changed, and that Claude Code's bundled Claude API skill can apply the swap with /claude-api migrate.

What it means for teams building AI products and agents

Cost. For routing, classification and extraction calls with short prompts, the per-token cut is large enough to outweigh the extra tokens from the new tokenizer, but recount real prompts with model set to claude-haiku-5-5 before you put a number in a budget. Agents that stuff long documents into one call are different: once a prompt passes 100,000 tokens the higher rate applies, so chunking, retrieval and prompt caching matter more than they did. Our guide to controlling AI agent running costs covers how to set those budgets.

Latency. Adaptive thinking adds time on requests where Haiku 4.5 ran without thinking. For tight latency paths, start at a low effort level and measure, rather than assuming Haiku 4.5 response speeds carry over.

Governance. Haiku 5.5 runs safety classifiers that can decline a request with stop_reason: "refusal", and there is no server-side fallback, so your client needs its own handling and, where it matters, a fallback model. Thinking blocks only work in the account that produced them, which affects multi-tenant services that replay stored conversations through a different account.

Platform gaps. Priority Tier is not supported on Haiku 5.5, so teams with a Priority Tier commitment on Haiku 4.5 need to plan capacity separately. On Bedrock, which lacks structured outputs, prefill replacements have to use tools.

When not to switch yet. Stay on Haiku 4.5 for now if you depend on Priority Tier, if your code relies on prefill or fixed sampling settings you cannot test quickly, or if most of your traffic sits in long prompts where the over-100,000 rate applies. Check the model deprecations page for how long Haiku 4.5 stays available, and plan the move as described in what to do when a model is deprecated. Run your evals on both models first; what AI agent evals catch explains what that should cover.

Frequently asked questions

What is Claude Haiku 5.5?

Claude Haiku 5.5 is Anthropic's small Claude model for high-volume, latency-sensitive work, announced on 7 October 2026. It has a 1M token context window, up to 128k output tokens and adaptive thinking with an effort setting. It is available on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry under the model ID claude-haiku-5-5.

Is Claude Haiku 5.5 cheaper than Haiku 4.5?

Per token, yes for prompts up to 100,000 tokens, where the list price is well below Haiku 4.5. Two things narrow the gap. The new tokenizer counts the same text as more tokens, and prompts over 100,000 tokens pay a higher rate. Recount your real prompts on Haiku 5.5 and price them with its own rates before you assume a saving.

Can I switch from Haiku 4.5 by changing the model ID?

Only if your requests avoid the five breaking changes. Manual thinking budgets, non-default temperature or top_p, assistant prefill, the older computer use tool and edited earlier turns with thinking blocks all return errors. Also check code that reads the first content block, and raise small max_tokens limits. Agents on Claude Managed Agents need only the model name changed.

Does Claude Haiku 5.5 refuse more requests?

Haiku 5.5 runs safety classifiers that can decline a request, returning a stop reason of refusal, and Anthropic offers no server-side fallback for it. That does not mean it refuses often, but your client must handle the case. For user-facing agents, log refusals, show a clear message and consider routing the request to another model where your policies allow it.

When should a team stay on Claude Haiku 4.5?

Stay put for now if you rely on Priority Tier, which Haiku 5.5 does not support, or if most calls use long prompts that would hit the higher rate above 100,000 tokens. Also wait if you cannot yet test code that uses prefill or fixed sampling settings. Check the deprecation date for Haiku 4.5 and plan the migration before it arrives.

Share this post

Contact us to build the right product

Talk to our engineers about your application, the systems it connects to, and what you want to build next.

Get in touch
Two people discussing work with a laptop