veasy

Claude · Sonnet 5.5 · AI models · Anthropic · Claude Code · AI agents · research

Claude Sonnet 5.5: my first read of the launch numbers

Anthropic released Claude Sonnet 5.5 on September 28, 2026. Same price as Sonnet 5, 30%+ faster, up to 30% cheaper per task. What the numbers say and what I will change in my own agent setup.

Sonnet 5.5 keeps Sonnet 5's price ($2 in / $10 out per million tokens) but uses fewer tokens per task, runs 30%+ faster, and jumps from 10.3% to 70.6% on Terminal-Bench 4.0. Opus 5.5 still leads on open-ended work that needs judgment.

Chart comparing Claude Sonnet 5 and Sonnet 5.5: Terminal-Bench 4.0 rises from 10.3% to 70.6%, OSWorld 2.1 from 57.0% to 80.1%, price unchanged at $2 / $10 per million input / output tokens, output 30%+ faster. Source: Anthropic.

Anthropic shipped Claude Sonnet 5.5 yesterday, September 28, 2026. It is the second model in the Claude 5.5 family, a week after Opus 5.5. I read the announcement and the benchmark table this morning. These are my notes, written from the point of view of someone who runs a small fleet of Claude Code agents every day and pays attention to the bill.

I have not moved my agents to it yet, so everything below comes from Anthropic's published numbers, not my own tests.

The short version

  • Same price as Sonnet 5. $2 per million input tokens, $10 per million output tokens, $0.20 per million for cache reads, $2.50 for cache writes.
  • Cheaper per task anyway. It needs fewer tokens to finish the same work. Anthropic measured up to 30% lower cost per task.
  • Faster. Output generation is 30%+ faster than Sonnet 5, the fastest Sonnet so far.
  • Much better at agentic coding. Terminal-Bench 4.0 goes from 10.3% (Sonnet 5) to 70.6%.
  • Opus 5.5 is still the model for hard, open-ended work. Anthropic says this plainly, and the benchmark gaps agree.

API model ID: claude-sonnet-5-5. It is available on the Claude Platform, AWS, Google Cloud and Microsoft Azure, with zero data retention like Opus 5.5.

The benchmark table

Benchmark

Sonnet 5.5

Sonnet 5

Opus 5.5

Terminal-Bench 4.0 (agentic coding)

70.6%

10.3%

66.4%

FrontierCode 1.1 (Main)

52.1% (Xhigh)

42.4%

54.4%

CursorBench 4.0

55.5%

34.1%

57.8%

GDPval-AA v2.1 (knowledge work, Elo)

1844

1449

1846

AA-Briefcase v1.1

1811

1359

1822

Humanity's Last Exam (with tools)

64.5%

54.9%

67.7%

OSWorld 2.1 (computer use)

80.1%

57.0%

81.8%

Chartography (no tools)

61.6%

15.6%

64.4%

Source: Anthropic's launch page. The Opus 5.5 Terminal-Bench score is at Xhigh effort, its best run.

Two rows stand out to me. Terminal-Bench is the one closest to how I use Claude: long command-line sessions where the model has to plan, run things, read the output and recover. Going from 10.3% to 70.6% moves Sonnet from "I would not trust it alone in a terminal" to "worth a real try". The other is GDPval-AA, a test of real work across 44 occupations: Sonnet 5.5 lands 2 points below Opus 5.5 and about 400 above Sonnet 5.

Cost per task matters more than price per token

The headline price did not move, so a quick look at the pricing page says nothing changed. The change is in how many tokens the model spends to get something done.

Anthropic plots score against cost per task at every effort level. On several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost. On FrontierCode at High effort it scores 10 points above Sonnet 5 at roughly one fifteenth of the cost per task.

The customer numbers point the same way. Balyasny reported about 121k tokens per answer on their finance tasks against 497k for Sonnet 5. Slack saw about 14% fewer output tokens on their Slackbot evals without touching their prompts. Lovable measured a third fewer tool calls in their coding evals.

For anyone running agents all day, that is the number to watch. My agents mostly do well-scoped work: fix this bug, update that config, check this deploy. That is exactly the category Anthropic says Sonnet 5.5 is built for.

Effort levels are now the main dial

Effort decides how long Claude thinks and how carefully it checks its work. The defaults differ by surface:

  • Claude Code and the Claude apps: Medium
  • Claude Platform (API): High

Higher effort is not always better. One footnote in the announcement is a nice example: on FrontierCode, Sonnet 5.5 scored lower at Max than at Xhigh. At Max it more often ran Claude Code's multi-agent code-review skill, and in two cases that led to a timeout or to edits outside the task's scope, which the benchmark penalizes. So more effort sometimes means more work nobody asked for.

A sensible starting point: run agents on Medium and raise effort only for tasks that actually fail at Medium.

Where Opus 5.5 still wins

Anthropic is direct about this: in their testing and external testing, Opus 5.5 is still clearly stronger at complex, open-ended work that needs sustained judgment. Benchmarks show one facet, and the gap on Humanity's Last Exam and FrontierCode is still there.

The split that makes sense to me: Opus 5.5 designs the architecture and makes the judgment calls, Sonnet 5.5 implements and handles the high-volume everyday tasks. One of the early testers, Creator, described using it exactly this way for game development.

Safety changes worth knowing

  • Cyber safeguards. Sonnet 5.5's cybersecurity capability is close to Opus 5's, so it is the first Sonnet launched with cyber safeguards. Normal bug fixing is unaffected. Higher-risk security tasks visibly fall back to Sonnet 5.
  • Biology safeguards are the same as Sonnet 5's.
  • Distillation protection. It is the first Sonnet with classifiers that block reasoning extraction, and it expands preserved thinking, so the model's thinking stays tied to the account that created it. If you switch accounts in the middle of a Claude Code session, read Anthropic's docs article on this change.
  • Alignment audit. Across roughly 1,850 scenarios, it matches or improves on Sonnet 5 on most measures. Opus 5.5 still scores slightly better overall.

Migration notes

  • Change the model string to claude-sonnet-5-5.
  • If you run Sonnet with thinking off, you need to switch to the new between_tools setting before moving to Sonnet 5.5. Anthropic's migration guide covers the details.
  • Price does not change, so no budget update is needed, but re-check your effort settings, since that is where the savings come from.
  • Claude Haiku 5.5, for high-volume and cost-sensitive work, is announced for the coming weeks.

How I would test it

The cheapest honest test: move one low-risk agent to Sonnet 5.5 at Medium effort for a week and compare its token use and failure rate against a week on Sonnet 5. Anthropic's numbers come from their evals and their early customers, and your workload is neither.

If you are running Claude agents too and you have tried Sonnet 5.5 already, leave a comment below. I am curious whether the "fewer tokens per task" claim holds up outside Anthropic's evals.

Source: Introducing Claude Sonnet 5.5, Anthropic

ShareFacebookLinkedInX

Comments