Sonnet 5.5: 30% Faster and Cheaper? API Changes Before Upgrading
What Anthropic says about Claude Sonnet 5.5: the conditions behind 30% faster and up to 30% lower cost per task, benchmarks, and API changes before upgrading.
Data checked: September 29, 2026. This article uses Anthropic's own primary sources only: the Sonnet 5.5 announcement, the Claude Platform docs overview for Sonnet 5.5, and What's new in Claude Sonnet 5.5. Every number described as "official" comes from the vendor's own tests and customer reports. None of it is our own measurement.
Anthropic released Claude Sonnet 5.5 on September 28, 2026. It calls the model a clear upgrade over Sonnet 5, with output generated 30%+ faster and up to 30% lower cost per task for most work. Here is the official launch video from the Claude channel:
The short version
- The price per token did not change. What changed is how many tokens a task needs. List pricing is the same as Sonnet 5 ($2 per million input tokens, $10 per million output tokens). "Up to 30% less" comes from finishing the same work with fewer tokens and fewer tool calls, not from a lower unit price. Whether your bill drops 30% depends on your own workload.
- Official benchmarks are far ahead of Sonnet 5. For example, Terminal-Bench 4.0: Sonnet 5.5 scores 70.6%, Sonnet 5 scores 10.3%. The effort settings are not identical across every row, listed below.
- Upgrading is more than changing the model ID. The official docs list five changes that break code running on Sonnet 5 (including how to turn thinking off, forced tool choice, and parameters such as
temperature). Test before you ship.
What Anthropic says about speed and cost
The announcement makes three separate claims. Do not read them as one:
- Speed: output is generated 30%+ faster than Sonnet 5, and Anthropic calls it its fastest Sonnet to date.
- Cost per task: up to 30% lower for most work, because the model typically needs fewer tokens to do the same job.
- Cost at a fixed benchmark score: on several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score at about a tenth of the cost per task, or less (Low effort on CursorBench 4.0, Medium effort on Terminal-Bench).
Claim 3 compares cost at a given score. It does not mean your everyday work becomes ten times cheaper. And the "up to" in claim 2 is a ceiling, not an average.
The announcement also quotes customer reports. Each company measured in its own environment, so the numbers are not comparable with each other:
| Reported by | Figure quoted in the announcement |
|---|---|
| Balyasny Asset Management | About 121k tokens per answer on 2,441 finance tasks, versus 497k for Sonnet 5 |
| Box | Results 2.4x faster with 12% fewer total tokens |
| Zendesk | Tickets processed 20% faster |
| Atlassian | Rovo Agents up to 30% faster |
| Slack | Better than Sonnet 5 on almost all offline Slackbot evals, in fewer steps |
| Lovable | Roughly half the shell runs to finish a task versus Sonnet 5 |
Official benchmarks (Sonnet 5.5 vs Sonnet 5 and Opus 5.5)
| Evaluation | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% |
| FrontierCode 1.1 Main (agentic coding) | 46.2% (Max) | 42.4% | 54.4% |
| CursorBench 4.0 (agentic coding) | 55.5% | 34.1% | 57.8% |
| GDPval-AA v2.1 (knowledge work, Elo) | 1844 | 1449 | 1846 |
| AA-Briefcase v1.1 (knowledge work, Elo) | 1811 | 1359 | 1822 |
| Humanity's Last Exam (with tools) | 64.5% | 54.9% | 67.7% |
| OSWorld 2.1 (computer use, partial) | 80.1% | 57.0% | 81.8% |
| Chartography (chart recognition, no tools) | 61.6% | 15.6% | 64.4% |
Read the table with Anthropic's own footnotes in mind. On Terminal-Bench, the Opus 5.5 result is at Xhigh effort. On FrontierCode, Sonnet 5.5 is at Max effort, and Anthropic says Max effort caused timeouts and extra edits under one code-review skill. GDPval-AA and AA-Briefcase were measured on a pre-release deployment with a structured-output bug that has since been fixed, and Anthropic expects a small understatement. In short, these are vendor-published results. Different effort levels and tool setups cannot be compared directly, and whether they carry over to your tasks is something you have to test.
How to think about cost (same unit price)
Pricing per million tokens: $2 input, $10 output, $0.2 cache read, $2.5 cache write, the same as Sonnet 5, with a 1M-token context window and 128K max output (official docs). The claude-sonnet-5.5 model in the BazaarLink catalog uses the same unit prices. For a side-by-side spec comparison with Sonnet 5, see Should you switch from Claude Sonnet 5 to 5.5?.
The "savings" have to be verified on your own bill. Pick 20 to 30 real tasks, run them on both Sonnet 5 and 5.5 with the same prompts, and record input/output tokens, retries and pass rate. Then compare the cost per completed task, not the unit price.
API changes before you upgrade (official docs)
The following are the changes listed in the official What's new document. They describe Anthropic API behavior. When you call through BazaarLink, test with your own real requests before going live.
- The way to turn thinking off changed. Sending
thinking: {"type": "disabled"}to Sonnet 5.5 returns a 400. To turn off up-front thinking, use{"type": "between_tools"}, and only at high effort or below (xhigh and max return a 400). - Forced tool use is not supported.
tool_choiceset toanyor to a specific tool returns a 400. Onlyauto(the default) andnonework. To get schema-valid tool input, the docs suggestautowith strict tool use, or structured outputs. - Thinking blocks are tied to the model and the conversation. Each block records which model produced it. Sonnet 5.5 can read Sonnet 5's thinking blocks but not Opus 5.5's, and no other model can read Sonnet 5.5's. If you switch models mid-conversation, the turns after the switch run without the earlier reasoning. On newer accounts, replaying a block after editing the system prompt, tools or earlier messages returns a 400, so keep conversations append-only.
- The old computer-use tool is not accepted.
computer_20251124returns a 400. Move tocomputer_toolset_20260801. - Advisor tool pairings changed. With Sonnet 5.5 as the executor, Opus 4.8, Opus 4.7 and Sonnet 5 cannot be advisors and return a 400.
Two more changes do not fail a request but change behavior:
- Text between tool calls now arrives in
thinkingblocks. With the defaultdisplay: "omitted"that text is empty. If your UI streams those notes to users, it goes quiet between tool calls with no error. - Setting
temperature,top_portop_kto a non-default value returns a 400.
Effort levels were also recalibrated: the same level does not produce the same amount of thinking as on Sonnet 5, and the docs recommend re-running your effort sweep instead of carrying settings over. Start at high in general. For agentic coding and multistep tool use, start at medium for well-specified tasks and move to high for harder ones. For chat and latency-sensitive work, start at medium or low.
Who should switch?
Worth trying first: everyday coding, bug fixes and producing documents or slides, the well-scoped work Anthropic positions it for, especially if speed or token usage matters to you.
Test carefully before switching if:
- your code uses
thinking: disabled, forcedtool_choice, customtemperature, or the old computer-use tool. - your interface shows the text between tool calls.
- conversations move between models or edit earlier messages.
FAQ
Is Sonnet 5.5 cheaper than Sonnet 5?
The unit price is the same. What Anthropic claims is that the model needs fewer tokens for the same work, so cost per task can fall by up to 30%. The real number depends on your workload.
How was "30% faster" measured?
The announcement only says output is generated 30%+ faster than Sonnet 5. The customer figures (for example Zendesk 20% faster, Atlassian up to 30% faster) come from each company's own tests. Actual latency depends on prompt length, output length and effort level.
Can I keep using Sonnet 5?
Yes. If you upgrade, go through the official migration checklist (the five items above are its breaking changes) and re-run your effort sweep.
Are the scores in this article our own tests?
No. All benchmark and customer figures are relayed from Anthropic's announcement, with Anthropic's condition notes kept.
Sources
- Anthropic: Introducing Claude Sonnet 5.5
- Claude Platform docs: Claude Sonnet 5.5, What's new in Claude Sonnet 5.5
- Official video: Introducing Claude Sonnet 5.5 (Claude YouTube channel)
Enter a model and token volume to estimate the cost of an API request.
Open the cost calculatorTWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API
