October 9, 2026

Claude Haiku 5.5: pricing, what breaks from Haiku 4.5 and when to move

Claude Haiku 5.5 lists at $0.10/$0.50 per million tokens, but budget_tokens now errors and text counts about 30% more tokens. Cost and migration plan.

News

Insight

Anthropic released Claude Haiku 5.5 on October 7, 2026, at $0.10 per million input tokens and $0.50 per million output tokens, against $1 and $5 for Haiku 4.5. The catch is in the docs, not the launch post: code written for Haiku 4.5 can return 400 errors on the new model, and the same text counts as about 30% more tokens. The date to watch is October 15, 2026, the earliest date Anthropic lists for retiring Haiku 4.5; any retirement also needs at least 60 days' notice, and none has been published.

Everything below comes from Anthropic's own pages. Where we calculate something, we say so and show the inputs.

What Anthropic launched on October 7

The model ID is claude-haiku-5-5, a fixed ID with no date suffix and no alias. On Amazon Bedrock it's anthropic.claude-haiku-5-5. Per Anthropic's Haiku 5.5 overview, it has a 1M-token context window and 128K max output (300K on the Batch API beta with the output-300k-2026-03-24 header). Thinking is adaptive, with effort defaulting to medium.

Anthropic's positioning, per that same overview, is "high-volume, latency-sensitive tasks such as classification, extraction, and routing." It's available on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. Anthropic lists the earliest retirement date for Haiku 5.5 as "Not sooner than October 7, 2027."

Haiku 5.5 pricing against Haiku 4.5 and Sonnet 5.5

All prices are per million tokens, from Anthropic's pricing page.

ModelInputOutputCache read
Haiku 5.5 (prompts up to 100,000 tokens)$0.10$0.50$0.01
Haiku 5.5 (prompts over 100,000 tokens)$0.50$2.50$0.05
Haiku 4.5$1$5$0.10
Sonnet 5.5$2$10$0.10

Haiku 5.5 is the only current model Anthropic's pricing page says is priced by prompt length. The page says prompt length "counts all of its input tokens, including cache reads and cache writes," so a long cached prefix can push a request into the higher tier. Anything that uses the 1M window pays the $0.50/$2.50 rate, five times the headline.

The batch rates are $0.05 input and $0.25 output up to 100k tokens, and $0.25 and $1.25 above. Haiku 4.5's batch rates are $0.50 and $2.50.

The effective price after the tokenizer change

List price isn't what you'll pay. Anthropic's migration guide says the same text produces "approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5," and that the exact increase depends on the content. Our calculation, from those two sources:

  • Up to 100k tokens: $0.10 x 1.3 = $0.13 input and $0.50 x 1.3 = $0.65 output, per million Haiku-4.5-equivalent tokens of text. Against $1 and $5, that's 87% lower on both.
  • Over 100k tokens: $0.50 x 1.3 = $0.65 input and $2.50 x 1.3 = $3.25 output. That's 35% lower than $1 and $5.

The 100k threshold is counted in Haiku 5.5 tokens, so a prompt of about 77k Haiku 4.5 tokens already crosses it, and the savings hold only as far as the 30% figure holds for your content.

Both figures leave out thinking. With adaptive thinking on by default, thinking tokens count toward max_tokens, and we found no figure for how many to expect. Our 87% is a ceiling for short prompts, not a forecast.

What breaks when you move from Haiku 4.5

The release notes put it bluntly: "Code written for Claude Haiku 4.5 can break on Claude Haiku 5.5." The details are in the Haiku 5.5 migration guide. Haiku 4.5's IDs are claude-haiku-4-5-20251001 and claude-haiku-4-5.

Thinking

{"type": "enabled", "budget_tokens": N} returns a 400. Use "thinking": { "type": "adaptive" } with "output_config": { "effort": "medium" }. Because adaptive thinking is on by default, a response can start with thinking blocks, so select content blocks by type rather than by position. Thinking blocks come back with an empty thinking field unless you send "display": "summarized".

Sampling

A non-default temperature or top_p, or any top_k, returns a 400. If you send them at all, temperature must be 1 and top_p 0.99. Sending both temperature and top_p also returns a 400.

Prefill

Per the guide, "Claude Haiku 5.5 rejects it with a 400 error, even with thinking turned off." Any classifier that forces an output format by pre-filling the assistant turn needs a rewrite.

Computer use and conversation state

computer_20250124 returns a 400 on the Claude API and Google Cloud. Use computer_toolset_20260801. Thinking blocks are account-bound, so replay a stored conversation only through the account that produced it. Separately, sending a thinking block back after changing system, tools or earlier messages returns a 400, so keep conversations append-only.

Changes that are easy to miss

Three changes are easy to miss. Refusals are real: "Claude Haiku 5.5 runs safety classifiers that can decline a request, and it has no server-side fallback," so your code has to handle stop_reason: "refusal" itself. Our piece on how Anthropic bills refusals before any output covers the fields involved, though that billing change is scoped to Fable 5.1, Fable 5, Opus 5.5 and Opus 5, and that page doesn't say how Haiku 5.5 refusals are billed. A max_tokens value tuned for Haiku 4.5 may cut output off, since the tokenizer counts more. And "Priority Tier is not supported on Claude Haiku 5.5," which matters if you hold a Priority Tier capacity commitment on Haiku 4.5.

Managed Agents users only need to change the model name. Everyone else can try the guide's own tool, the /claude-api migrate command in Claude Code.

When Haiku 4.5 stops working

Anthropic's deprecations page lists claude-haiku-4-5-20251001 as Active, with a tentative retirement of "Not sooner than October 15, 2026." That's a floor. We read no retirement notice for it as of today, and Anthropic promises "at least 60 days' notice before model retirement," counted from the notice, not from October 15. A working Haiku 4.5 integration isn't about to stop. Amazon Bedrock and Google Cloud set their own dates.

If you want a model for how that timeline plays out, the Sonnet 4.5 retirement is the live example, with a hard date of November 30, 2026.

Sonnet 5.5 cache reads and the Max and Team credits

Anthropic also cut Sonnet 5.5's cache-read price on October 7, from $0.20 to $0.10 per million tokens, or 0.05x the base input price instead of 0.1x. The API release notes add: "Cache writes and all other prices are unchanged." Our earlier Sonnet 5.5 migration write-up shows the old $0.20; the figure in the table above is current.

On the plan side, Anthropic's launch announcement says Max 5x users get $100 in credits per month, Max 20x users $200, and Team subscribers "up to $500, pooled across their users." The credit is for use on the Claude Platform. The post says it rolls out "this week," so it wasn't confirmed live on launch day. We haven't read terms on rollover, expiry, or Pro and Enterprise plans, so we aren't stating any. If you're a startup instead, there's a separate route: Claude for Startups credits and seats.

For Anthropic's benchmark claims, the launch page reports its own numbers: OSWorld 2.1 (offline subset) 72.4% against 15.7% for Haiku 4.5, Terminal-Bench 4.0 39.2% against 0.0%, and Humanity's Last Exam without tools 45.9% against 10.2%. No independent lab has reproduced them that we've seen. Anthropic's own line: "Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0."

What this means for your AI budget and migration plan

The decision is which high-volume calls (classification, extraction, routing, subagents) move to Haiku 5.5, and when. Nothing forces you before October 15, and any retirement needs at least 60 days' notice, so there's no deadline until a notice appears.

Our judgment, not a sourced fact: replay a sample of real production traffic on claude-haiku-5-5 before switching anything. Re-baseline token counts and max_tokens against the roughly 30% increase, and measure thinking tokens yourself, because we found no published figure. Check whether any workload crosses 100k prompt tokens, counted with the Haiku 5.5 tokenizer (about 77k Haiku 4.5 tokens), where the saving shrinks to 35% on our numbers. Skip Haiku 5.5 for anything on Priority Tier. Keep Sonnet 5.5 on agentic coding, as Anthropic itself advises. Treat the Max and Team credit as a one-off offset until its terms are published.

For the seat side of the same budget, see AI coding tools cost per developer vs. hiring one. For other releases and price changes, see AI model releases and pricing: what changes for engineering teams.

FAQ

How much does Claude Haiku 5.5 cost?

For prompts up to 100,000 tokens, $0.10 per million input tokens and $0.50 per million output tokens. Above 100,000 tokens it's $0.50 and $2.50. Batch halves the lower tier to $0.05 and $0.25. Because text counts about 30% more tokens than on Haiku 4.5, our calculation puts the effective price near $0.13 and $0.65 per Haiku-4.5-equivalent million tokens, before thinking tokens.

Will my Haiku 4.5 code work on Haiku 5.5?

Not unchanged. budget_tokens, non-default temperature, top_p or top_k, assistant prefill and computer_20250124 (on the Claude API and Google Cloud) each return a 400. Adaptive thinking is on by default, so responses can start with thinking blocks. Refusals have no server-side fallback.

Is Claude Haiku 4.5 being retired?

No retirement notice was published in what we read. The deprecations page lists claude-haiku-4-5-20251001 as Active with a tentative retirement of "Not sooner than October 15, 2026," and Anthropic commits to at least 60 days' notice.

Is Haiku 5.5 better than Sonnet 5.5 for coding agents?

Anthropic says no. It states that Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0." Haiku 5.5 is positioned for classification, extraction and routing.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now