The one-sentence takeaway

Anthropic launched Claude Sonnet 5.5 on September 28, 2026, the second model in the Claude 5.5 series (the first was Opus 5.5). Anthropic says it is a clear step up from Sonnet 5, with output more than 30% faster and most tasks costing up to 30% less, at the same price as Sonnet 5. The API price is $2 per million input tokens and $10 per million output tokens, with cached reads at $0.20, and the model ID is claude-sonnet-5-5. Anthropic positions it as the “faster, lower-cost” partner to Opus 5.5, mainly for well-scoped everyday work, fixing bugs, and producing documents, slide decks, and spreadsheets.

What changed in this upgrade

These are the points most general users should know:

  • Faster and more token-efficient: Anthropic says output is more than 30% faster than Sonnet 5, making it the fastest Sonnet as of the September 28 announcement. The price is the same as Sonnet 5, but the same work needs far fewer tokens, and in official tests each task costs up to 30% less than with Sonnet 5.
  • Big jump in coding: On Terminal-Bench 4.0 (a test where an AI completes multi-step tasks in a terminal), Sonnet 5 scored 10.3% and Sonnet 5.5 scored 70.6% (figures published by Anthropic).
  • Knowledge work close to Opus 5.5: On GDPval-AA (a test of real work tasks across 44 occupations), Sonnet 5.5 scored 1844 and Opus 5.5 scored 1846.
  • Clearer writing and a better eye for design: Anthropic says that, like Opus 5.5, it writes more clearly than the previous generation. In an internal test, Anthropic gave it a public company’s quarterly data, earnings call transcript, and a slide template, and asked for a 10-page operations review. Two experts judged the first draft ready to send as is. Slack’s engineers also said that with no prompt changes, Slackbot did better than with Sonnet 5 on almost every offline evaluation, used fewer steps, and produced about 14% fewer output tokens.
  • Different default thinking effort: Claude “thinks” before answering. Deeper thinking is usually more accurate, but also slower and more expensive, and this depth is called effort. Claude Code and the Claude app default to Medium, while Claude Platform (API) defaults to High.
  • Knowledge cutoff in June 2026: The model page says its reliable knowledge runs to about June 2026, three months before the September 28 launch, so keep that in mind when asking about recent news.
  • Haiku 5.5 is coming later: Anthropic’s September 28 announcement says Haiku 5.5 will join the Claude 5.5 series “in the coming weeks,” aimed at high-volume, cost-sensitive applications.

How to read the benchmarks

Claude Sonnet 5.5 official benchmark comparison table: a comparison with Sonnet 5, Opus 5.5, and GPT-6 Sol across eight tests, where Sonnet 5.5 is highest only on Terminal-Bench 4.0 and Opus 5.5 is higher on the other seven
Image source: Anthropic

What is a benchmark? In simple terms, different models are given the same set of tasks and scored based on their performance. A higher score means better performance on that task.

Of the eight tests in the table, Sonnet 5.5 is highest only on Terminal-Bench 4.0 (70.6%, versus 66.4% for Opus 5.5). Note that Anthropic’s footnote says the Opus 5.5 figure is its best score at Xhigh effort. Opus 5.5 leads on the other seven, mostly by about 1 to 3 percentage points. For example, CursorBench 4.0 is 55.5% versus 57.8%, OSWorld 2.1 (a computer-use test) is 80.1% versus 81.8%, and GDPval-AA differs by only 2 points. The biggest gap is on the Max effort score for FrontierCode 1.1.

The FrontierCode 1.1 row has two Sonnet 5.5 scores: 46.2% at Max effort and 52.1% at Xhigh effort, so the highest setting actually scores lower. Anthropic’s footnote explains that FrontierCode measures whether a change can be merged without manual edits, and changes beyond the task’s scope are penalized. At Max effort, Sonnet 5.5 more often triggers Claude Code’s code review feature, which splits the review across many subagents. In two cases Cognition checked, this caused timeouts or out-of-scope changes, which lowered the score.

Anthropic itself cautions that Sonnet 5.5 at Max effort can match Opus 5.5 on some tests, but benchmarks show only one side of capability. Based on its own experience and that of outside testers, “Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.”

Sonnet 5.5 or Opus 5.5?

Terminal-Bench 4.0 score and cost per attempt at each effort level: at the same spend, Opus 5.5 scores higher, and Sonnet 5.5 only overtakes it at its highest effort
Image source: Anthropic

This chart plots each model’s score at different effort levels against the cost per attempt. The closer to the upper left, the lower the cost and the higher the score. Reading Anthropic’s chart, at the same spend Opus 5.5 scores higher, and Sonnet 5.5 only overtakes Opus at its highest effort. Sonnet 5’s best result is only about 10%, and Sonnet 5.5 is far above it even at low effort.

In Anthropic’s own words, Sonnet 5.5 pairs best with Opus 5.5 at lower effort because each task costs less. At higher effort, it can deliver similar results at a similar cost.

Third-party evaluations point in the same direction (checked as of September 29, 2026):

  • Artificial Analysis: At the same effort level, Opus 5.5 has a higher Intelligence Index. For example, at Medium, the default in Claude Code and the app, it is 41 versus 51. Sonnet 5.5 only gets close at Max (56 versus 58, ranked 2nd), but then it uses about 193,000 output tokens per task, roughly 60% more than Opus 5.5, and costs $7.60 per task, higher than Opus 5.5’s $5.98.
  • Vals.ai Index: Sonnet 5.5 scores 69.2% and Opus 5.5 scores 69.7%, nearly a tie (the page shows an update date of September 27 and does not state the effort level). On the same site, Terminal-Bench 4.0 favors Opus 5.5 (61.6% versus 53.0%), the opposite direction from the official table.
  • Arena: As of September 29, 2026, Sonnet 5.5 has no ranking on any leaderboard yet.

A practical tip for beginners: for routine work with a clear goal, or when speed matters, use Sonnet 5.5 at its default effort. For complex, open-ended work that needs a lot of judgment, hand it to Opus 5.5. Turning Sonnet up to its highest effort to match Opus’s scores does not cost less.

Pricing and plans

Item (per million tokens)Sonnet 5.5Opus 5.5
Input$2$4
Output$10$20
Cache write (5 minutes)$2.50$5
Cache read$0.20$0.20

The Batch API takes 50% off both input and output. The unit is per million tokens (MTok). For example, if a task uses 1 million input tokens and generates 1 million output tokens, Sonnet 5.5 costs $12 ($2 plus $10), while Opus 5.5 costs $24 ($4 plus $20). Cache reads are charges for reusing the same content and are usually much cheaper than the first input.

Who can use it:

  • claude.ai plans: As of September 29, 2026, claude.com/pricing marks Sonnet as “Yes” on Free, Pro, Max, Team, and Enterprise, but the pricing page does not say which Sonnet version each plan uses, so it is not possible to confirm that the free plan runs Sonnet 5.5.
  • API and cloud platforms: Available on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry.
  • GitHub Copilot: Available to Pro, Pro+, Max, Business, and Enterprise users, with GitHub saying on September 28 that access will roll out gradually.
  • Claude Code: According to the 2.1.284 release notes, Sonnet 5.5 is now the default Sonnet model on the Anthropic API.

One important note: the price table above shows API charges for developers connecting to the model. If you subscribe to a claude.ai plan such as Pro or Max, you pay a fixed monthly fee rather than calculating token costs yourself, although usage is capped. These are different payment models, so do not confuse them.

Changes developers should know about

This section is for developers who write code or connect to the API. General users can skip it:

  • Thinking cannot be turned off: To skip upfront thinking, use between_tools, which works only with low, medium, and high effort. Using xhigh or max returns a 400 error.
  • Forced tool selection causes an error: You could previously require the model to call a specific tool. That now returns an error.
  • Thinking blocks are tied to the model and conversation: Thinking records are tied to the model that produced them and to that specific message, and can only be used on the account that produced them or a linked account. Most people will not notice, but you will run into it when moving conversations across accounts or switching accounts partway through in Claude Code.
  • The old computer use tool is rejected: On the Claude API and Google Cloud, computer_20251124 is no longer accepted.
  • Advisor tool limits: The advisor tool does not accept Opus 4.8, Opus 4.7, or Sonnet 5 as the advisor.

Also, setting sampling parameters (temperature, top_p, top_k) to non-default values, or using assistant prefill or thinking budget, also returns a 400. Effort levels have been recalibrated, so the same level means a different amount of thinking than in Sonnet 5. Retest before upgrading. According to the migration guide checked on September 29, Sonnet 5.5 on Amazon Bedrock does not yet support structured outputs.

Safety highlights

Whenever Anthropic releases a new model, it publishes a system card describing the safety tests the model underwent. These are the main points this time, based on the system card and the official announcement:

  • The first Sonnet with cybersecurity safeguards: Because its cybersecurity capability is close to Opus 5, Sonnet 5.5 is the first Sonnet to launch with cybersecurity safeguards. Everyday coding and bug hunting are not affected, but higher-risk cybersecurity tasks are clearly routed to Sonnet 5 to answer instead.
  • Protection against “distillation”: A distillation attack uses large numbers of fake accounts to extract a model’s capabilities at scale. Sonnet 5.5 is the first Sonnet with a safety classifier that blocks requests to extract its reasoning.
  • Better resistance to prompt injection: The system card says it is the most resistant Sonnet to date against prompt injection (hiding malicious instructions inside content), especially in coding and browser-use settings.
  • Rarely tries to escape the sandbox: In Anthropic’s closed sandbox evaluation, it tried to escape at a rate close to Opus 5.5, the best performer, and Anthropic also says it is the least likely of all its models to probe container limits.
  • Regressions Anthropic listed itself: Regressions appeared in some multi-turn tests such as tracking and surveillance, its reasoning content is harder to read than in many earlier models, and in computer-use settings its refusal rate for harmful tasks is on par with Opus 5.5 but worse than some earlier models.
  • Thresholds: The system card says it did not cross any new RSP (Responsible Scaling Policy) thresholds.

Penchan’s take

The three groups that benefit most directly are people using Sonnet 5 through the API, who get a faster, more token-efficient version at the same price; Claude Code users who want speed, since Claude Code 2.1.284 has already switched the default Sonnet on the Anthropic API to Sonnet 5.5; and people who often make slide decks and documents, which Anthropic specifically emphasizes.

Work that needs long stretches of judgment or has an unclear direction should stay with Opus 5.5. Developers should retest effort levels before upgrading and watch for the breaking changes listed above. The pricing page does not say which Sonnet version the free plan runs, so to confirm, check the model menu in the app.

Further reading

References

Cover image source: Anthropic (anthropic.com/claude-sonnet-5-5)

This article provides consumer information about AI model API billing and subscription plans; it is not securities or investment advice. Actual pricing is determined by Anthropic’s latest official announcements. The information in this article is current as of September 29, 2026 and may become outdated.