Claude Sonnet 5.5 Ranks #2 at 56, Behind Only Opus 5.5

Anthropic released Claude Sonnet 5.5 on September 28, 2026, and independent benchmarker Artificial Analysis ranks it second on its Intelligence Index with a score of 56, two points behind Claude Opus 5.5 and ahead of both GPT-6 Astra and Claude Fable 5.1. The mid-tier model keeps Sonnet 5’s price of $2 / $10 per million input / output tokens. Anthropic says it runs more than 30% faster and costs up to 30% less per task. At maximum effort, though, Artificial Analysis measured the highest token use it has ever recorded.

Intermediate

Claude Sonnet 5.5 wordmark over a view of Earth seen through a spacecraft window
Image credit: Anthropic

What Anthropic Reports

Sonnet 5.5 is the second model in the Claude 5.5 family, a week after Opus 5.5. Anthropic presents it as the faster, cheaper complement to Opus, strongest at “well-scoped everyday tasks,” bug fixing and producing documents, slides and spreadsheets. On Anthropic’s own evaluations it comes close to Opus 5.5 across most categories and passes it on one:

  • Terminal-Bench 4.0 (agentic coding): 70.6%, against 66.4% for Opus 5.5 and 10.3% for Sonnet 5
  • CursorBench 4.0: 55.5% (Opus 5.5: 57.8%; Sonnet 5: 34.1%)
  • OSWorld 2.1 (computer use): 80.1% (Opus 5.5: 81.8%; Sonnet 5: 57.0%)
  • GDPval-AA v2.1 (knowledge work): 1,844 (Opus 5.5: 1,846; Sonnet 5: 1,449)
  • Humanity’s Last Exam: 64.5% (Opus 5.5: 67.7%; Sonnet 5: 54.9%)
  • Chartography (chart understanding, no tools): 61.6% (Opus 5.5: 64.4%; Sonnet 5: 15.6%)

Anthropic credits the lower cost per task to fewer tokens and fewer tool calls, not to a price change. Several early-access customers quoted in the announcement report similar results. Box says the model was “more accurate, 2.4x faster, and used 12% fewer total tokens” than its predecessor, and Lovable saw about a third fewer tool calls on coding jobs. The model has five effort levels (low, medium, high, xhigh and max). According to Unite.AI, the default is medium in Claude Code and the Claude apps, and high on the API.

What Artificial Analysis Measured

Bar chart of the Artificial Analysis Intelligence Index v4.3.2 with Claude Opus 5.5 at 58, Claude Sonnet 5.5 at 56, Claude Fable 5.1 and GPT-6 Astra at 53, and Claude Sonnet 5 at 38, above a scatter plot of index score against cost per task
Intelligence Index v4.3.2 scores (top) and score versus cost per task (bottom). Image credit: Artificial Analysis

At max effort, Sonnet 5.5 scores 56, an 18-point jump over Sonnet 5 (38). Artificial Analysis finds it level with Opus 5.5 on its knowledge-work evaluations: 1,811 against 1,822 Elo on AA-Briefcase, and 1,844 against 1,846 on GDPval-AA. The gaps are on Humanity’s Last Exam and SciCode, where it trails by about six points, and on factual accuracy (AA-Omniscience), where it scores 54% to Opus 5.5’s 66%. Lower effort settings score much less: 47 at high, 41 at medium and 36 at low.

Bar charts of Terminal-Bench 4.0, with Claude Sonnet 5.5 first at 63.6% ahead of Claude Opus 5.5 at 59.6% and GPT-6 Astra at 59.1%, and Terminal-Bench-Science 0.1, with Sonnet 5.5 third at 53.3%
Terminal-Bench 4.0 and Terminal-Bench-Science scores. Image credit: Artificial Analysis

Artificial Analysis’s own Terminal-Bench 4.0 run puts Sonnet 5.5 first at 63.6%, ahead of Opus 5.5 (59.6%) and GPT-6 Astra (59.1%), and up from 14.1% for Sonnet 5. Anthropic reports 70.6% on the same benchmark. The two figures come from different test harnesses and should not be compared directly.

The top score has a cost. Artificial Analysis says Sonnet 5.5 at max effort used “~193k Output Tokens per Intelligence Index Task. This is the highest token use we have measured, around 60% higher than Opus 5.5 (max).” That is roughly seven times GPT-6 Astra at max. The resulting cost is about $7.60 per index task, around 50% more than Sonnet 5, even though per-token prices did not change.

Safeguards

Sonnet 5.5 is the first Sonnet model to ship with the cyber safeguards Anthropic had previously applied only to its top-tier models. Anthropic rates its cybersecurity capability as comparable to Opus 5.5. High-risk cyber requests fall back to Sonnet 5, and Anthropic’s Cyber Verification Program offers tiered access for security work. Unite.AI describes a three-stage classifier pipeline with “99.43% recall on the cyber harm coverage set,” and notes that users should “expect increased refusals even on benign cybersecurity tasks.” It is also the first Sonnet with classifiers that block extraction of its reasoning, which is a defence against distillation, and preserved thinking is tied to the account that generated it.

What This Means

The two claims about cost are not necessarily contradictory, because they measure different things. Anthropic’s “up to 30% less per task” comes from its own testing against Sonnet 5. Artificial Analysis’s $7.60 figure is for max effort specifically, where Sonnet 5.5 spends more tokens than any model the benchmarker has tested. At max effort, Sonnet 5.5 comes within two points of Opus 5.5 while costing somewhat more per index task than Opus does. At high effort it scores 47, close to GPT-6 Sol at about the same cost per task.

For researchers and course builders on a budget, the effort setting now matters about as much as which model you pick. The same Sonnet 5.5 spans roughly 20 index points between low and max effort. The benchmark to watch is the one run at the setting you will actually deploy. Claude Haiku 5.5, which Anthropic describes as built for high-volume, cost-sensitive work, is due “in the coming weeks.”

Related Coverage

This post was drafted with AI assistance and reviewed by RITS staff.

Sources