Claude Sonnet 4.5 Deprecation: Migrating to Sonnet 5.5

Anthropic announced the deprecation of Claude Sonnet 4.5 with retirement set for November 30, 2026. Learn how to update API calls and migrate to Sonnet 5.5.

By ailog · Published Oct 1, 2026 · 6 min read · How we work
In short
  • Anthropic deprecated Claude Sonnet 4.5 on September 30, 2026, with API retirement set for November 30, 2026.
  • Claude Sonnet 5.5 serves as the designated replacement, sharing the exact same rate structure: $2 per 1M input tokens, $0.2 per 1M cached input tokens, and $10 per 1M output tokens.
  • Requests lacking an explicit thinking block run adaptive thinking automatically on Sonnet 5.5 instead of turning thinking off.
  • Setting thinking to disabled produces a 400 error on Sonnet 5.5; workloads requiring no up-front reasoning must supply between_tools instead.
  • Pre-filling assistant turns, specifying thinking token budgets, and supplying non-default sampling arguments trigger immediate 400 client errors.

Deprecation Schedule and Impacted Endpoints

On September 30, 2026, Anthropic published an official update to its platform release notes announcing the formal deprecation of claude-sonnet-4-5-20250929. The model has entered the deprecated stage of the provider lifecycle, meaning it remains operable temporarily but is no longer advised for current deployments. The planned shutdown date across the primary Claude API is November 30, 2026. Once this retirement date arrives, any inbound API transactions directed to the identifier will terminate in a failure response.

This deprecation applies directly to workloads managed on Anthropic-operated environments, specifically the direct Claude API, Claude Platform on AWS, and Microsoft Foundry. External cloud environments such as Amazon Bedrock and Google Cloud maintain independent lifecycle calendars and may establish different operational cutoffs for hosted model snapshots. Anthropic provides at least 60 days’ notice prior to retiring publicly accessible models, giving engineering teams an explicit window to adjust production codebases.

Teams running workloads through Claude Managed Agents face minimal disruption: switching the model string over to claude-sonnet-5-5 completes the transition. In contrast, applications that issue raw requests to the Messages API require code auditing. Sonnet 5.5 alters default reasoning mechanisms and removes legacy parameters that previously functioned on Sonnet 4.5.

Thinking Mode Differences and Response Handling

The operational boundary between Claude Sonnet 4.5 and Claude Sonnet 5.5 centers on how the models handle internal reasoning tokens. On claude-sonnet-4-5-20250929, sending a request without a thinking payload left reasoning disabled. On claude-sonnet-5-5, requests omitting a thinking property activate adaptive thinking automatically. If an engineering team attempts to turn off internal thinking using {"type": "disabled"}, the Sonnet 5.5 endpoint rejects the payload with a 400 invalid_request_error.

To bypass reasoning before generating an answer on Claude Sonnet 5.5, clients must configure the thinking type as between_tools. This setting skips upfront deliberation while preserving brief reasoning blocks between iterative tool interactions. Furthermore, the display attribute defaults to omitted on Sonnet 5.5, returning empty text within reasoning blocks alongside a signature. To receive readable intermediate notes similar to the default behavior of Sonnet 4.5, callers must explicitly request display: "summarized".

Model Name Default Thinking (No Field Sent) Accepted thinking.type Values Default display
Claude Sonnet 5.5 On (adaptive) “adaptive”, “between_tools” “omitted”
Claude Sonnet 5 On (adaptive) “adaptive”, “disabled” “omitted”
Claude Sonnet 4.6 Off “adaptive”, “disabled”, “enabled” (deprecated) “summarized”
Claude Sonnet 4.5 Off “disabled”, “enabled” “summarized”
Claude Haiku 4.5 Off “disabled”, “enabled” “summarized”

Downstream parser code must also adapt. On Sonnet 4.5, applications frequently assumed that textual output resided in the first element of the response array (content[0].text). Because Sonnet 5.5 routinely prefixes generation with reasoning structures, code relying on static array indexing will break when it encounters a non-text block. Callers must iterate over elements or filter by block.type == "text". Additionally, when running tool calling loops, callers must preserve and return all intermediate reasoning blocks to the API without modifications.

Breaking Parameters and Payload Constraints

Migrating raw requests from Sonnet 4.5 to Sonnet 5.5 involves removing several configurations that prompt instant 400 client status errors on the newer architecture. The API enforces strict validation rules to eliminate deprecated settings:

  • Disabled thinking: Submitting thinking type disabled returns a 400 error. Callers must either adopt adaptive thinking or configure between_tools.
  • Thinking token budgets: Explicit numerical limits such as budget_tokens trigger a 400 error on Sonnet 5.5.
  • Sampling arguments: Setting non-default values for temperature, top_p, or top_k causes a 400 error on Sonnet 5.5, while official Python SDK releases starting at v1.0 raise a local TypeError.
  • Assistant prefill: Supplying prefilled assistant text in the message history is rejected on Sonnet 5.5.
  • Forced tool choice: Forcing specific tool execution modes that bypass model selection logic results in an invalid request.
  • Effort boundaries: Specifying between_tools alongside xhigh or max effort generates an error; effort must be restricted to low, medium, or high.

Another operational constraint involves conversation consistency when executing under between_tools. Applications cannot vary output_config.effort across turns within an ongoing multi-message exchange; modifying the effort parameter mid-stream returns a 400 error. For systems requiring dynamic per-turn effort calibration, requests must operate under standard adaptive thinking instead.

Updating Request Syntax and Automation Tools

Constructing valid requests for Sonnet 5.5 requires restructuring configuration blocks. The standard pattern replaces sampling overrides with output_config, allowing callers to specify the desired effort level directly. The following cURL example illustrates a properly formed payload utilizing medium effort without deprecated parameters:

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-5-5",
    "max_tokens": 4096,
    "messages": [{
      "role": "user",
      "content": "Analyze the trade-offs between microservices and monolithic architectures"
    }],
    "output_config": {
      "effort": "medium"
    }
  }'

For workflows migrating away from upfront thinking, the configuration requires specifying between_tools inside the thinking object alongside an accepted effort level like high:

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-5-5",
    "max_tokens": 16000,
    "thinking": {"type": "between_tools"},
    "output_config": {"effort": "high"},
    "messages": [{"role": "user", "content": "..."}]
  }'

Developers operating inside Claude Code can automate these code modifications across repository files using the bundled Claude API skill. Running /claude-api migrate this project to claude-sonnet-5-5 scans project directories, updates model strings, eliminates invalid sampling values, addresses assistant prefill occurrences, and adapts thinking parameters. The tool requests directory confirmation before applying changes and supports adjustments for Amazon Bedrock and Claude Platform on AWS configurations.

Pricing Structure and Model Cost Profiles

Moving from Claude Sonnet 4.5 to Claude Sonnet 5.5 does not increase base token unit rates. Sonnet 5.5 maintains standard Sonnet pricing: $2 per 1M input tokens, $0.2 per 1M cached input tokens, and $10 per 1M output tokens. However, because adaptive thinking activates by default, total expenditures may climb if applications generate unmanaged internal reasoning tokens. Internal thinking tokens are billed directly as standard output tokens at $10 per 1M units. Development teams should review their token consumption profiles using /tools/llm-api-cost-calculator/ to anticipate overall budgetary shifts.

Model Identifier Input Cost (per 1M) Cached Input (per 1M) Output Cost (per 1M)
claude-sonnet-5-5 $2 $0.2 $10
claude-opus-5-5 $4 $0.2 $20
claude-fable-5.1 $10 $0.25 $50
claude-haiku-4-5 $1 $0.1 $5
gpt-6.1-sol $2 $0.1 $10
gemini-3.1-pro-preview $2 $0.2 $12

Because max_tokens applies cumulatively across thinking tokens and regular generated text, applications operating with constrained output budgets risk truncating user-facing replies if internal reasoning consumes the allocation. Teams running cost-sensitive batch jobs or evaluating tier alternatives can explore options through /posts/choosing-an-llm-api-model-by-cost-tier/ and reference baseline rates on /tools/llm-api-pricing/.

Migration Verification Checklist

To ensure applications successfully migrate away from claude-sonnet-4-5-20250929 before the November 30, 2026 retirement date, teams should execute the following verification steps:

  • Audit usage: Visit the Usage page in Claude Console, export the account CSV log, and locate all API keys issuing requests to claude-sonnet-4-5-20250929.
  • Update model IDs: Replace instances of claude-sonnet-4-5-20250929 with claude-sonnet-5-5 across configuration files and agent runners.
  • Eliminate sampling settings: Strip temperature, top_p, and top_k from request construction logic to prevent 400 response errors or SDK TypeErrors.
  • Remove prefill and forced tools: Check prompts to ensure no trailing assistant roles are supplied for prefill, and confirm tool configurations do not enforce forced choice.
  • Configure thinking mode: Choose between default adaptive reasoning or thinking type between_tools, ensuring thinking disabled is never passed.
  • Adjust effort constraints: Verify that output_config.effort is restricted to low, medium, or high when using between_tools, and keep effort consistent across turns.
  • Update response parsing: Ensure output deserialization reads blocks by type instead of accessing static indices like content[0].text.
  • Preserve multi-turn blocks: Pass thinking blocks back into message arrays unaltered during consecutive tool execution cycles.

Completing these checks ensures uninterrupted service across primary API pipelines and guarantees full operational compatibility with Claude Sonnet 5.5 before legacy endpoints cease operation.

Sources
  1. Anthropic Claude Platform release notes — 2026-09-30
  2. Migrating to Claude Sonnet 5.5
  3. Model deprecations

Related