You have built the Claude API into internal AI apps and business systems, with claude-sonnet-5 as the default model. Then Sonnet 5.5 comes out at the same price, and you are told, "If just changing the ID makes it faster and cheaper, let's switch right away." Meanwhile, you also have not decided which of the workloads running on Opus 5.5 can be moved down to Sonnet 5.5.
Sonnet 5.5 keeps the same pricing, but some settings in requests written for Sonnet 5 will return 400 errors if sent as is. Some changes also alter only what appears on screen without causing errors. Within what we could confirm in the official documentation, we sort out the items to check before switching.
Same pricing as Sonnet 5, and the ID is claude-sonnet-5-5 with no date
Sonnet 5.5 was released on September 28, 2026, and is the second model in the Claude 5.5 family. Anthropic positions it as "a faster, lower-cost model that complements Opus 5.5." Among the current models listed on the models overview page, Sonnet 5.5 and the models above and below it compare as follows.
| Item | Sonnet 5.5 | Opus 5.5 | Haiku 4.5 |
|---|---|---|---|
| Claude API ID | claude-sonnet-5-5 | claude-opus-5-5 | claude-haiku-4-5-20251001 |
| Input / output (per million tokens) | $2 / $10 | $4 / $20 | $1 / $5 |
| Cache reads | $0.20 | $0.20 | $0.10 |
| Context window / max output | 1M / 128K | 1M / 128K | 200K / 64K |
| Thinking | Adaptive (on by default) | Adaptive (always on) | Extended thinking |
| Default effort in the API | high | medium | Not supported |
| Retirement no earlier than | September 28, 2027 or later | September 22, 2027 or later | October 15, 2026 or later |
Sonnet 5.5 pricing is the same as Sonnet 5, including prompt caching and batch processing ($2.50 for 5-minute cache writes, $4 for 1-hour cache writes, and 50% off both input and output with the Batch API). The tokenizer is also the same as Sonnet 5, so the same text produces the same number of tokens.
Anthropic's announcement page says output is more than 30% faster than Sonnet 5, and that in its own tests the cost per task was up to 30% lower (because fewer tokens are used). Your own workloads will not necessarily see the same ratios.
It is available on the Claude API, Amazon Bedrock (ID anthropic.claude-sonnet-5-5), Claude Platform on AWS, Google Cloud and Microsoft Foundry. On every platform other than Amazon Bedrock, the ID is claude-sonnet-5-5.
Five places where Sonnet 5 code can return 400 errors
The official "What’s new in Claude Sonnet 5.5" lists five breaking changes that affect code running on Sonnet 5.
| Setting sent to Sonnet 5 | Handling in Sonnet 5.5 | Replacement Destination |
|---|---|---|
thinking: {"type": "disabled"} | 400 error | {"type": "between_tools"} (effort at high or below) |
any / tool in tool_choice | 400 error (also when counting tokens) | auto + strict tool use, with the prompt specifying when to call the tool |
| Editing earlier history and resending thinking blocks | 400 error under certain conditions | Make the conversation append-only |
computer_20251124 (Claude API, Google Cloud) | 400 error | computer_toolset_20260801 |
| Opus 4.8, Opus 4.7 or Sonnet 5 as the advisor | 400 error | An advisor that Sonnet 5.5 accepts |
The first row is the change also singled out in the "Getting started" section of the announcement page. In Sonnet 5.5, adaptive thinking runs even if nothing is specified. Workloads that turned thinking off should be rewritten to use between_tools, the lowest setting. It cannot be combined with xhigh or max effort. For requests that do not use tools, the response will be text only.
# Sonnet 5
client.messages.create(model="claude-sonnet-5", thinking={"type": "disabled"}, ...)
# Sonnet 5.5
client.messages.create(model="claude-sonnet-5-5", thinking={"type": "between_tools"},
output_config={"effort": "high"}, ...)
The second row applies to workloads that force a specific tool call for classification or field extraction. With auto, the model may answer without calling a tool, so specify in the prompt when it should call one. On Amazon Bedrock, structured outputs (including strict tool use) are not available for Sonnet 5.5, so the guidance is to validate tool inputs in your code.
The third row is a check that applies by default on the Claude API, Amazon Bedrock and Google Cloud for accounts created on or after August 31, 2026 (UTC). If you modify the system prompt, tools or earlier messages and then resend Sonnet 5.5 thinking blocks, you get an error. To change instructions, use a system message inserted partway through the conversation.
There is one more point where the shape of responses changes without any error. Notes longer than one or two sentences written between tool calls are, by default, returned as thinking blocks with empty contents. Screens that stream intermediate progress to users go silent without any error. Using between_tools or specifying thinking.display restores the previous behavior.
Behavior that changes without errors
Even without changing your code, the following differences appear.
- The effort scale has changed. The same
highdoes not produce the same amount of thinking as on Sonnet 5, and the official docs ask you to re-measure. As a starting point, use high for general workloads, medium for agentic workloads with well-defined steps, and medium or low for latency-sensitive chat. - More kinds of requests are refused. Refusals are returned as
stop_reason: "refusal", andstop_detailscontains one of five categories: cyber, bio, frontier_llm, reasoning_extraction or general_harms. The server-side fallback in the Claude API (beta) retries only cyber and frontier_llm refusals on Sonnet 5. - The minimum cache length is shorter. The minimum length of a cacheable prompt has dropped from 1,024 tokens on Sonnet 5 to 512 tokens.
- Thinking blocks can be used only by the account that created them. If sent from a different account, the blocks are dropped and the request succeeds.
If you are moving from Sonnet 4.6 or earlier, there are even more changes. Setting non-default values for temperature, top_p or top_k returns a 400 error, and thinking budgets (budget_tokens) are not available either. The same text also uses about 30% more tokens than on Sonnet 4.6, Sonnet 4.5 and Haiku 4.5.
Workloads to move down from Opus 5.5 and up from Haiku 4.5
We covered Opus 5.5 pricing and how to review costs in A cheaper new model does not automatically lower your bill. Here we look at where Sonnet 5.5 fits.
The official "Choosing a model" page positions Opus 5.5 for complex agentic coding and enterprise work, and Sonnet 5.5 for everyday coding and business workloads where you want both speed and capability. The announcement page lists, for Sonnet 5.5, Sonnet 5 and Opus 5.5 respectively, 70.6%, 10.3% and 66.4% on Terminal-Bench 4.0 (the Opus 5.5 figure is at xhigh effort) and 55.5%, 34.1% and 57.8% on CursorBench 4.0. The same page also says Opus 5.5 is clearly stronger on complex work involving long chains of judgment, so the closeness of the numbers alone cannot decide which workloads to move down.
When routing workloads from Opus 5.5 to Sonnet 5.5, the following points change.
- Sonnet 5.5 does not read thinking blocks from Opus 5, Opus 5.5, Fable or Mythos. If you switch from Opus 5.5 partway through a conversation, subsequent responses do not carry over that thinking (the request succeeds, and dropped blocks are not billed).
- Fast mode (research preview) is available for Opus 5.5, Opus 5 and Opus 4.8, but not for Sonnet 5.5.
- Opus 5.5 cannot turn thinking off, but Sonnet 5.5 can turn off upfront thinking with
between_tools.
When moving up from Haiku 4.5, the unit price rises and the same text also uses more tokens, so the migration guide asks you to re-measure costs. The announcement page says Claude Haiku 5.5, aimed at high-volume workloads, is planned to arrive within a few weeks. If you run volume-driven workloads on Haiku 4.5, the editorial team's view is to wait for Haiku 5.5 and compare before deciding whether to move up to Sonnet 5.5. We also cover how to assign models to each workload in How to decide which model to use for development.
Sonnet 5 remains available until at least June 30, 2027
The "Model deprecations" page lists claude-sonnet-5 as Active and states that it will not be retired before June 30, 2027. Customers using a public model are notified at least 60 days before it is retired.
It is actually the older generations whose deadlines are closer. claude-sonnet-4-5-20250929 may be retired on or after September 29, 2026, and claude-haiku-4-5-20251001 on or after October 15, 2026. Neither has been marked Deprecated yet. These dates mean "not retired before this date"; the actual retirement date is set by announcement. These dates apply to the Claude API, Claude Platform on AWS and Microsoft Foundry; Amazon Bedrock and Google Cloud set their own deadlines. You can check which models are used by which keys in the CSV exported from the Usage page in the Claude Console.
Steps for switching

The diagram maps the items to check to the official pages for verifying them. From here on, these are the editorial team's recommendations based on the official documentation.
- Find the relevant settings in your code. Search for
"disabled",tool_choice,temperature,top_p,top_k,computer_20251124, advisor settings, and code that rewrites history and resends it, then compare them against the table above. The migration guide also describes using/claude-api migratein Claude Code to handle the rewrite and build a checklist. - Compare two or three effort levels. Do not reuse the value you used on Sonnet 5; measure quality and cost per request with the same kinds of inputs as production.
- Check screens that show intermediate progress. Look at the actual screens to see whether nothing is shown between tool calls anymore.
- Check how refusals are handled. Decide whether receiving
stop_reason: "refusal"should stop processing as an error or retry with another model. - Make the model ID a configuration value and switch per workload. Do not switch all workloads at once; switch them one at a time and keep the ability to revert to Sonnet 5.
If model names are scattered across multiple places in your code, step 5 turns into a refactoring project. We covered the idea of consolidating access into a single entry point in Should you put a gateway in front of your internal AI?.
On September 30, 2026, we cross-checked by directly opening the Claude release notes (the September 28, 2026 entry), the Claude API release notes, the models overview in Claude Platform Docs, the Sonnet 5.5 model page, What’s new, migration guide and prompting guide, pricing, Model deprecations, Choosing a model, and Anthropic's Sonnet 5.5 announcement page. Because we had no API key, we did not send requests to Sonnet 5.5 and did not verify error contents, speed, token counts, cost per request, or behavior on Amazon Bedrock, Google Cloud or Microsoft Foundry. Benchmark figures are Anthropic's published values.
For switching models in internal systems built on the Claude API, or reworking them into an architecture where models are easy to swap, contact GleamHub.
Sources
- Release notes — Claude Help Center
- Claude Platform release notes — Claude Platform Docs
- Models overview — Claude Platform Docs
- Claude Sonnet 5.5 — Claude Platform Docs
- What’s new in Claude Sonnet 5.5 — Claude Platform Docs
- Migrating to Claude Sonnet 5.5 — Claude Platform Docs
- Prompting Claude Sonnet 5.5 — Claude Platform Docs
- Pricing — Claude Platform Docs
- Model deprecations — Claude Platform Docs
- Choosing the right model — Claude Platform Docs
- Introducing Claude Sonnet 5.5 — Anthropic









