Haiku 5.5 costs $0.10/$0.50 per MTok vs $1/$5 on Haiku 4.5, but a tokenizer change and a 5x tier above 100K tokens decide your real bill.
Anthropic released Claude Haiku 5.5 on October 7 at $0.10 per million input tokens and $0.50 per million output tokens, one tenth of Claude Haiku 4.5's $1 and $5. Two details decide whether your bill really falls by 90%: a newer tokenizer that Anthropic says produces roughly 30% more tokens for the same text, and a price tier that jumps fivefold once a prompt passes 100,000 tokens.
What happened
Anthropic's release notes list claude-haiku-5-5 as available on October 7 on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. It has a 1M-token context window and up to 128K output tokens on the synchronous API. Anthropic positions the model for high-volume, latency-sensitive work such as classification, extraction and routing, and says it will not be retired before October 7, 2027.
Pricing, from Anthropic's pricing page (per million tokens):
| Haiku 5.5, prompts up to 100K | Haiku 5.5, prompts over 100K | Haiku 4.5 | |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output | $0.50 | $2.50 | $5.00 |
| Cache read | $0.01 | $0.05 | $0.10 |
| 5-minute cache write | $0.125 | $0.625 | $1.25 |
| Batch input / output | $0.05 / $0.25 | $0.25 / $1.25 | $0.50 / $2.50 |
Claude Code v2.1.293, also released October 7, added the model and made it the default Haiku model on the Anthropic API. That is the default for the Haiku slot, not the default model overall. GitHub also made Haiku 5.5 generally available in Copilot the same day, for Pro, Pro+, Max, Business and Enterprise plans, billed "at provider list pricing under usage-based billing" according to GitHub's changelog.
Why it matters
At list price, the cut is large even after the tokenizer change. Anthropic's Haiku 5.5 what's new page and migration guide say the same input text produces "approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5. The exact increase depends on the content." The 30% below is that figure; measure your own traffic before budgeting on it.
The longer-context rule is the part to plan around. Claude 4.6 and later models, except Haiku 5.5, are priced flat across the 1M window; Haiku 5.5 is not. A prompt over 100,000 tokens pays $0.50 input and $2.50 output instead of $0.10 and $0.50. The pricing page does not spell out whether the higher rate covers the whole request or only the tokens past 100K, but it says "a prompt of over 100,000 tokens pays higher prices", and Anthropic's Haiku 5.5 overview lists the output rate as "$2.50 / MTok for prompts over 100,000 tokens". Output rate depending on prompt length points to the whole request, so this article assumes that. Treat the figures below as the conservative case until Anthropic confirms.
Technical details
The calculations below use list prices only. "Same workload" means the job that took 50,000 input and 2,000 output tokens on Haiku 4.5, which under a 30% inflation becomes 65,000 and 2,600 tokens on Haiku 5.5. All dollar figures are computed by ThreatFrontier from the published rates, not quoted from Anthropic.
| Request | Haiku 4.5 | Haiku 5.5 | Change |
|---|---|---|---|
| 50K in / 2K out, equal token counts | $0.0600 | $0.0060 | −90% |
| Same workload, 5.5 counts 30% more tokens (65K / 2.6K) | $0.0600 | $0.0078 | −87% |
| 80K in / 2K out on 4.5, which is 104K / 2.6K on 5.5 | $0.0900 | $0.0585 | −35% |
| 120K in / 2K out, equal token counts | $0.1300 | $0.0650 | −50% |
The cliff. If 5.5 counts 30% more tokens, a prompt that is 76,923 tokens on Haiku 4.5 becomes 100,000 tokens on Haiku 5.5, the last token billed at the low tier. At that size the request costs about $0.0113 on Haiku 5.5 with 2.6K output tokens, against about $0.0869 on Haiku 4.5. One token later, under the whole-request reading, it costs about $0.0565, five times as much, because the input rate moves from $0.10 to $0.50 for every token in the prompt.
Input cost of one prompt by size. Haiku 5.5 jumps 5x at 76,923 Haiku 4.5 tokens (100K Haiku 5.5 tokens; Anthropic says ~30% more tokens, approximate) but stays below Haiku 4.5.
Haiku 5.5 stays cheaper than Haiku 4.5 on both sides of the line. Above 100K, 1.3 times the tokens at $0.50 works out to an effective $0.65 per Haiku 4.5-equivalent million input tokens against $1.00, a 35% saving, and the same holds for output at $3.25 against $5.00. The cliff is a budgeting problem, not a reason to stay on the older model: a workload that sits at 70K–90K Haiku 4.5 tokens will see its cost per request jump about fivefold between prompt sizes while still ending up cheaper than before.
Caching moves the numbers. Cache reads cost $0.01 per million tokens up to 100K and $0.05 above it, so a large, stable prefix that is cached stays inexpensive (we covered caching and cost-per-task tuning on Sonnet 5.5). 5-minute cache writes rise fivefold too, from $0.125 to $0.625 per MTok.
Two API behaviours change migration work. Adaptive thinking is on by default, so responses can begin with thinking blocks. Setting the manual budget_tokens parameter returns a 400 error. Code that sends it to Haiku 4.5 must drop it (our Sonnet 4.5 migration guide covers the same pattern), and parsers must tolerate thinking blocks at the start of the response.
What defenders should do
This is a cost and migration story rather than a vulnerability, but teams running agent and classification pipelines have decisions to make:
- Re-measure tokens, don't scale by 30%. Run representative prompts through the token counter for
claude-haiku-5-5and compare against Haiku 4.5 counts. The inflation depends on content. - Find prompts near 100K. Log input token counts per request and flag anything between roughly 75K and 100K Haiku 5.5 tokens. Trimming context or caching a stable prefix is worth more here than anywhere else in the range.
- Set spend alerts on the long-context tier. A pipeline that quietly drifts over 100K prompts will see its unit cost rise fivefold with no code change.
- Remove
budget_tokensand handle thinking blocks before switching model IDs, and test with output length caps, since thinking output is billed as output. - Claude Code users: the default Haiku model on the Anthropic API is now Haiku 5.5 as of v2.1.293. If you pin models for cost control, check your configuration; our Claude Code token economics piece has the cost background.
- Copilot users: GitHub says Haiku 5.5 is billed at provider list pricing, but did not state a credit multiplier in the changelog. For how credits map to plans, see our explainer on Copilot Pro versus Pro+ AI credits.
What is still unclear
- Whole-request or marginal tier. Anthropic's pages imply, but do not state outright, that the over-100K rates apply to the entire request. This article assumes they do.
- Your tokenizer ratio. Anthropic says the exact increase depends on content, so 30% is a planning figure, not a guarantee.
- Copilot cost. GitHub's changelog states list-price billing but no credit multiplier; we did not find one.
- Quality claims. GitHub says early testing showed Haiku 5.5 matching Claude Sonnet 5 on many coding tasks with fewer tokens and steps. That is GitHub's statement, not an independent benchmark, and we have not verified it.
Sources
- Anthropic, Pricing
- Anthropic, Release notes
- Anthropic, Models overview
- Anthropic, Haiku 5.5 overview, What's new and migration guide
- Claude Code, Changelog
- GitHub, Claude Haiku 5.5 in GitHub Copilot