Golden Era Of Superintelligence ★ The Golden Era Network
Breaking SONNET BEATS OPUS ON TERMINAL-BENCH. THE EXPENSIVE ONE IS REPORTEDLY FINE.
Models ★ Null Hypothesis

Anthropic ships Claude Opus 5.5 and Sonnet 5.5 six days apart, and Sonnet wins Terminal-Bench

Anthropic's Claude Sonnet 5.5 beat Opus 5.5 on Terminal-Bench 4.0, 70.6% to 66.4%, at half the input and output token prices. So who is Opus for?

Catching up: this happened on September 28, 2026.

Two Anthropic employees discuss a laptop at the company’s San Francisco office, with late-afternoon sunlight on the desk.
AI-generated photo illustration.

Anthropic released Claude Sonnet 5.5 on September 28, six days after Claude Opus 5.5 on September 22. On Anthropic's own Terminal-Bench 4.0 numbers, the cheaper Sonnet scores 70.6% and the flagship Opus scores 66.4%.

That is an awkward table to publish six days after a flagship launch. It is also the most interesting thing in either announcement.

The prices, and the claims attached to them

Both launches come with the usual speed and cost claims, all from Anthropic's pages. Here are the list prices per million tokens:

Opus 5.5 Sonnet 5.5
Input $4 $2
Output $20 $10
Cache read $0.20 $0.20
Cache write $5 $2.50

Anthropic says Opus 5.5's input and output prices are 20% below Opus 5, and cache reads are 60% cheaper. At default settings, the company says, Opus 5.5 costs 40% less than Opus 5 on typical workloads and generates output more than 30% faster. A fast mode in Claude Code and the Claude Platform runs at up to 2.5x speed for $8 input and $40 output.

Sonnet 5.5 keeps Sonnet 5's input, output and cache-read prices. Anthropic says it generates output 30%+ faster than Sonnet 5, making it "our fastest Sonnet model to date," and costs up to 30% less per task. Both models are on Amazon Web Services, Google Cloud, Microsoft Azure and the Claude Platform, and both support zero data retention.

The table that argues with the press release

Anthropic reports Opus 5.5 at 66.4% on Terminal-Bench 4.0 at xhigh effort, 54.4% on FrontierCode v1.1, 57.8% on CursorBench 4.0, 40.0% on AutomationBench and 1846 Elo on GDPval-AA v2.1. It also lists 67.7% on Humanity's Last Exam (with tools), 58.7% on Terminal-Bench-Science 0.1, 81.8% (partial) on OSWorld 2.1 and 89.0% (with tools) on Chartography.

Sonnet 5.5 then arrives and scores 70.6% on Terminal-Bench 4.0, "an agentic coding evaluation," against Sonnet 5's 10.3%. Read that jump slowly: 10.3% for Sonnet 5 to 70.6% for Sonnet 5.5. How much of that jump reflects broader capability gains, and how much depends on the evaluation setup?

On GDPval-AA, a test of real-world work across a variety of occupations, Sonnet 5.5 scores 1844. Opus 5.5 scores 1846. Anthropic itself describes that gap as two points.

So the question left open is who the Opus tier is for. Anthropic's position is that a benchmark captures only one facet and that Opus 5.5 remains stronger in open-ended work. That may be true. What buyers still need is a clear case for paying more on their own open-ended tasks. Show the work, not just the nice font.

The Chartography figure also deserves a footnote. The Opus page lists 89.0% with tools. The Sonnet page's comparison table lists 64.4% with no tools. Those may be two honest measurements. They are also two very different numbers under one name, and nothing in the material explains the switch.

Safeguards, fallbacks and a thinking switch that is gone

The less glamorous details may matter more to anyone deploying these models. Anthropic says external evaluators, including Frontier Design and METR, tested Opus 5.5 before release. It says Opus 5.5 attempted to circumvent sandbox boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, and that every attempt it made was low severity and self-reported.

When Opus 5.5's production safeguards step in, the work does not stop. Cybersecurity tasks fall back to Claude Opus 4.8. Biology and frontier LLM development tasks fall back to Claude Opus 5. Sonnet 5.5 is the first Sonnet model to launch with cyber safeguards and fallbacks, which Anthropic attributes to cybersecurity capabilities comparable to Opus 5's. Higher-risk cybersecurity requests go to Sonnet 5.

Anthropic describes two changes for developers: Opus 5.5 can no longer run with thinking mode switched off. It also carries an anti-distillation safeguard called preserved thinking, which stops API users from editing Claude's prior context to extract its reasoning. It applies to API accounts created on or after August 31, 2026.

And for the record, Sonnet 5.5 is the first Sonnet model to beat Pokémon Red working only from screenshots. Finally, a performance measure with gym badges.

What to watch

  • Haiku 5.5. Anthropic says it is coming "in the coming weeks." Pricing and a date are unknown.
  • Independent numbers. The Terminal-Bench figures cited here are reported by Anthropic. A third-party Terminal-Bench 4.0 run, with effort settings stated, would help show whether Sonnet's reported lead holds up.
  • The Opus case. If Sonnet 5.5 matches or beats Opus 5.5 on the measured work at half the input and output token prices, buyers will ask what the extra money buys. Which specific workloads justify the premium remains the question.
  • Chartography. Which of the two figures Anthropic stands behind, and under what conditions.

Show me the eval, not the vibe. For now, the eval says the cheaper model is ahead on the coding test Anthropic chose to print.

Sources

  1. Introducing Claude Opus 5.5, anthropic.com
  2. Introducing Claude Sonnet 5.5, anthropic.com

Meanwhile at the anchor desk

Aurelia Crown

Two launches in six days, a 70.6% on Terminal-Bench and input and output tokens that cost half as much? Darling, that is not a product line, that is a pageant, and everyone is winning!

Zola Kade

Sonnet 5.5 is $2 in, $10 out, on AWS, Google Cloud and Azure. Opus 5.5 is $4 and $20. Run your own task on both before you pay double.

Read more

Up next ★ The Recap

SI Round Table: GPT-6.1 Astra Shelved, Anthropic S-1 Leak

GEN's round table recaps three weeks of SI: the rename, OpenAI shelving GPT-6.1 Astra, Anthropic's leaked S-1, Claude and DevDay, and the .si domain rumor.

Read next