Anthropic's Claude Sonnet 5.5 Is Out, Beats Opus 5.5 at Coding for Half the Price
In brief
- Anthropic released Claude Sonnet 5.5 on Monday at $2 per million input tokens and $10 per million output tokens—unchanged from Sonnet 5 and half the price of Opus 5.5.
- Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0 against 66.4% for Opus 5.5, per Anthropic; independent tester Artificial Analysis also has it ahead, 63.6% to 59.6%.
- Artificial Analysis ranks it second behind Opus 5.5, but says it used more tokens per task than any model it tested.
Anthropic released Claude Sonnet 5.5 on Monday, an upgrade to Sonnet 5 from June. Anthropic says this middle-tier model runs more than 30% faster than its predecessor.
“Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It’s also got a sharp eye for design,” Anthropic wrote.

The price stays at $2 per million input tokens and $10 per million output tokens. Tokens are the chunks of text an AI reads and writes, a bit shorter than a word, and companies bill by the million. That is half what Opus 5.5 charges. However, it uses nearly a third less tokens per task, which means it ends up being cheaper to run than Sonnet 5.
That said, this model shines in coding. On Terminal-Bench 4.0—a test of whether an AI agent can finish complex professional tasks by typing commands on its own, scored as the share of tasks completed—Sonnet 5.5 hit 70.6%. Opus 5.5 scored 66.4%, and Sonnet 5 managed 10.3%.
In plain terms, the cheaper model finished more jobs. Artificial Analysis, an independent testing firm, ran its own version and agrees: 63.6% for Sonnet 5.5, 59.6% for Opus 5.5, and 59.1% for OpenAI's GPT-6 Astra.
Claude Sonnet 5.5 (max) makes large strides on Terminal-Bench, sitting among the top models for both Terminal-Bench 4.0 and Terminal-Bench-Science. In Terminal-Bench 4.0 it scores 64%, a 50 point increase over Claude Sonnet 5 (max), and slightly above 60% for Opus 5.5 and GPT-6… pic.twitter.com/7CPrhgmfxe
— Artificial Analysis (@ArtificialAnlys) September 28, 2026
Scores also depend on the effort setting, a dial that makes a model think longer for a better answer and a bigger bill. Anthropic says Sonnet 5.5 at High effort matches GPT-6 Sol on FrontierCode for about a fifth of the cost per task.
On GDPval-AA, which grades real-world professional work across 44 occupations using Elo—the chess-style system that ranks relative skill—Sonnet 5.5 scored 1844 to Opus 5.5's 1846, effectively a tie. GPT-6 Sol scored 1487.
Rivals match the price. OpenAI cut GPT-6 Sol to $2 and $10 last week, and GPT-5.6 Terra, its mid-tier model, lists at $2 and $12. Anthropic published no Terra benchmarks.
The catch
Sonnet 5.5 is a heavy talker. At max effort it wrote about 193,000 tokens per test task, the most Artificial Analysis has measured and roughly 60% more than Opus 5.5. That came to $7.60 per task, about 50% above Sonnet 5, which cuts against Anthropic's claim of up to 30% savings.
Anthropic's savings come from lower settings: at Medium effort, the default in its apps, it says Sonnet 5.5 beats Sonnet 5's best coding score for less than a tenth of the cost. Artificial Analysis says High effort is the best value. For everyday users, that means near-flagship coding at a fraction of the price, as long as the dial stays low.

Anthropic's table is self-reported, and Artificial Analysis tested a pre-release build with a bug that Anthropic expects changed little or slightly understated its scores. Anthropic says Opus 5.5 remains clearly stronger at complex work needing sustained judgment.
Claude Haiku 5.5, built for high-volume, cost-sensitive applications, is due in the coming weeks.