Sonnet 4.5 stops answering on Nov 30, 2026. Sonnet 5.5 is a third cheaper per token, but prefill, forced tool use and thinking budgets now return 400 errors.
Anthropic has put a date on Claude Sonnet 4.5. On September 30, 2026 it deprecated claude-sonnet-4-5-20250929, and on November 30, 2026 the model retires on the Claude API. After that date, requests to it fail. That gives teams still pinned to Sonnet 4.5 61 days from the notice, just over Anthropic's stated minimum of 60 days' notice for publicly released models.
The recommended replacement is claude-sonnet-5-5, which shipped two days earlier, on September 28. Moving is not a one-line change. Sonnet 5.5 rejects several request shapes that Sonnet 4.5 accepted, and each of them comes back as a 400 error, not a degraded answer. Code that swaps only the model string will fail on its first request if it uses prefill, forced tool use, a thinking budget or a non-default temperature.
Who is affected
- Anyone calling
claude-sonnet-4-5-20250929on Anthropic-operated platforms: the Claude API, Claude Platform on AWS and Microsoft Foundry. Anthropic says those are the platforms its dates apply to. - Amazon Bedrock and Google Cloud users are on the partners' own retirement schedules, which can differ. Check the model tables for your platform rather than assuming November 30.
- Teams that do not know whether they still use it. Anthropic's suggested audit is the Console's Usage page: click Export and read the CSV, which breaks usage down by API key and by model.
The price goes down, the token count goes up
On the rate card, the move is a cut of a third:
| Per million tokens | Sonnet 4.5 | Sonnet 5.5 |
|---|---|---|
| Input | $3 | $2 |
| Output | $15 | $10 |
| Cache hit | $0.30 | $0.20 |
| 5-minute cache write | $3.75 | $2.50 |
| Batch input / output | $1.50 / $7.50 | $1 / $5 |
List prices fall by a third on every line. The new tokenizer and default thinking take back part of the saving.
Two things eat into that saving.
- The tokenizer. Sonnet 5.5 uses Sonnet 5's tokenizer, which Anthropic says produces about 30% more tokens for the same text than Sonnet 4.5 does. On plain text, $2 a million tokens times 1.3 is about $2.60 for the same words, against $3 before. That is our arithmetic, not an Anthropic figure, and the real ratio depends on your content.
- Thinking runs by default. On Sonnet 4.5, a request with no
thinkingfield ran without thinking. On Sonnet 5.5, the same request runs adaptive thinking, and thinking tokens bill as output. A pipeline that never thought before now pays for it unless it opts out.
Images cost more too. Sonnet 5.5 uses the high-resolution tier, up to 2,576 pixels on the long edge and 4,784 tokens per image. Sonnet 4.5 stopped at 1,568 pixels and 1,568 tokens. Anthropic's example is a 2000×1500 image, which costs about 2.5 times as many tokens on Sonnet 5.5.
What you gain is room: a 1M-token context window with no beta header, and 128K tokens of output.
Six requests that now return a 400
Each request shape on the left returns a 400 on Sonnet 5.5. The computer-use row is for the Claude API and Google Cloud; on Amazon Bedrock the target is computer_20251124.
1. Thinking settings
Sonnet 4.5 accepted thinking.type values of disabled and enabled. Sonnet 5.5 accepts only adaptive and between_tools.
{"type": "enabled", "budget_tokens": N}returns a 400. Remove the budget and set an effort level inoutput_config.effort. There is no fixed mapping from a budget to an effort level, so Anthropic suggests evaluating at two or three levels.{"type": "disabled"}also returns a 400. To keep running without up-front thinking, sendbetween_tools. It works only atlow,mediumorhigheffort; atxhighormaxit is a 400 as well.
# Sonnet 4.5
client.messages.create(
model="claude-sonnet-4-5-20250929",
max_tokens=16000,
thinking={"type": "enabled", "budget_tokens": 10000},
messages=[{"role": "user", "content": "..."}],
)
# Sonnet 5.5
client.messages.create(
model="claude-sonnet-5-5",
max_tokens=16000,
output_config={"effort": "high"},
messages=[{"role": "user", "content": "..."}],
)
Sonnet 4.5 had no effort parameter at all, so set one explicitly. The levels are low, medium, high, xhigh and max; the API default is high. Anthropic suggests medium as the starting point for well-specified agentic coding and medium or low for chat and other latency-sensitive work.
2. Forced tool use
A tool_choice of type any or tool returns a 400 on Sonnet 5.5, including on the token-counting endpoint. Send tool_choice: {"type": "auto"}, mark the tool strict: true so its input matches the schema, and say in the prompt when to call it, because the model can now answer without calling the tool. Strict tools need additionalProperties: false on every object, and a request can carry at most 20 of them. On Amazon Bedrock, structured outputs are not available for Sonnet 5.5, so send auto without strict and validate the tool input in your own code.
3. Assistant prefill
Sonnet 4.5 let you end a conversation with a partial assistant turn. Sonnet 5.5 rejects it: the conversation must end with a user message. Anthropic's replacements depend on what the prefill was for: structured outputs for a format, a system-prompt instruction to skip preambles, and moving a continuation into the user message.
4. Sampling parameters
A non-default temperature, top_p or top_k returns a 400. Remove them. If you use the Python SDK v1.0 or later, those parameters are gone from the request type and raise a TypeError before the request is sent.
5. The old computer use tool
Sonnet 5.5 does not accept computer_20250124, the version Sonnet 4.5 used, on any platform. On the Claude API and Google Cloud, move to the computer_toolset_20260801 toolset. On Amazon Bedrock, move to computer_20251124. If you send the fine-grained-tool-streaming-2025-05-14 beta header, remove it; alongside a toolset entry it is a 400. Set eager_input_streaming: true on each tool that needs it instead.
6. The old structured-output parameter
output_format is deprecated. Without the structured-outputs-2025-11-13 beta header it returns a 400. Move it to output_config.format.
Changes that break your parser, not your request
These return 200 and still break code.
- Responses can start with thinking blocks. Code that reads
content[0].textbreaks. Read content blocks bytype. Thinking text is omitted by default; setdisplay: "summarized"if you show it. - Text between tool calls moves. Notes longer than a sentence or two that the model writes between tool calls now arrive as
thinkingblocks, empty at the default display. An interface that streamed those notes goes quiet. - Thinking blocks are bound to the account and the conversation. Sonnet 5.5 blocks work only in the account that produced them or one linked to it; another account's blocks are dropped silently. For accounts created on or after August 31, 2026, replaying a block after editing earlier history is a 400, so keep conversations append-only.
- Tool-call escaping can differ. Parse tool
inputwith a standard JSON parser. - Beta headers. Remove any context-window header and
interleaved-thinking-2025-05-14. - Prompt caching gets cheaper to start. The minimum cacheable prompt drops from 1,024 tokens on Sonnet 4.5 to 512.
Refusals: new safeguards, and some are billed again
Sonnet 5.5 brings real-time cyber safeguards that code written for Sonnet 4.5 never met. A decline comes back as stop_reason: "refusal" with a stop_details.category: cyber, bio, frontier_llm, reasoning_extraction or general_harms. Security teams doing legitimate offensive work can apply to Anthropic's Cyber Verification Program.
Since September 24, 2026, refusals that arrive before any output are billed again when the category is bio, frontier_llm or reasoning_extraction, at the rates of the model that ran. Early refusals in other categories are still free, and mid-stream refusals were already billed. Server-side fallback (beta, Claude API only) retries cyber and frontier_llm declines on Sonnet 5; it does not retry the others.
What to do before November 30
- Find every caller. Export the Console usage CSV and grep your code and configs for
claude-sonnet-4-5. - Decide the target. Sonnet 5.5 is the recommended replacement, with retirement not sooner than September 28, 2027. Sonnet 4.6 is still active (not sooner than February 17, 2027) and accepts temperature and the old thinking settings, but it costs $3/$15 like Sonnet 4.5, rejects prefill too, and only postpones the work.
- Fix the six 400s above, then the parser changes.
- Re-baseline cost. Recount tokens with the token-counting endpoint, revisit
max_tokens(it now covers thinking plus text), and run your evaluations at two or three effort levels. - Handle refusals. Branch on
stop_reason: "refusal"and logstop_details.category. - Watch the next dates. Claude Haiku 4.5 retires not sooner than October 15, 2026, and Claude Opus 4.5 not sooner than November 24, 2026.
In Claude Code, /claude-api migrate this project to claude-sonnet-5-5 runs Anthropic's bundled migration skill. It swaps the model id, applies the breaking parameter changes, replaces prefills and calibrates effort, then leaves a checklist of what to verify by hand. It asks for the scope before it edits anything.