Global foundation-model progress briefing
English Edition中文
Enter keywords to search ingested stories.

Today / Thursday, August 13, 2026

limbo logolimbo

Data updated

Jun 30, 06:46 PM

Live sources

17

Ingestion status

Live ingest

ResearchThe Decoder

Anthropic's new Claude Sonnet 5 closes the gap to the pricier Opus model series

Summary

Anthropic released Claude Sonnet 5, which beats its predecessor Sonnet 4.6 across all benchmarks and even edges past the larger Opus 4.8 on the GDPval-AA v2 knowledge work test with a score of 1,618.

Original Article

Captured source content or English translation, normalized into this reading format.

Read Source

Ad

Skip to content

Anthropic's new Claude Sonnet 5 closes the gap to Opus model series

![Matthias Bastian](https://the-decoder.com/author/matthias-bastian/ "View all posts by Matthias Bastian")

Matthias Bastian[View the LinkedIn Profile of Matthias Bastian](https://www.linkedin.com/in/matthias-bastian-128b71b1/ "View the LinkedIn Profile of Matthias Bastian")

Jun 30, 2026

Image description

Anthropic

Key Points

  • Anthropic released Claude Sonnet 5, which the company calls its most agentic Sonnet yet. It can build plans on its own and use tools like browsers and terminals.
  • In benchmarks, Sonnet 5 beats its predecessor, Sonnet 4.6, across the board and closes in on the larger Opus 4.8. On real-world knowledge work tasks, it even edges past Opus 4.8.
  • The model is available now on all Anthropic platforms at an introductory discount, with pricing rising to standard Sonnet rates after August 2026.

Ask about this article…Search

Topics

Anthropic released Claude Sonnet 5. In benchmarks, it closes in on the larger Opus 4.8 and even beats it in some areas. The model is available now at an introductory price.

Anthropic calls it the most agentic Sonnet yet: it can build plans, grab tools like browsers and terminals, and work on its own at a level that just months ago only bigger, pricier models could pull off, according to the company. Sonnet 5 is meant to close that gap.

Benchmarks show a clear jump over Sonnet 4.6

Anthropic's published benchmarks show Sonnet 5 beating its predecessor Sonnet 4.6 in every tested category while gaining ground on the pricier Opus 4.8. On agentic coding, Sonnet 5 hits 63.2 percent on SWE-bench Pro, up from 58.1 percent for Sonnet 4.6. Opus 4.8 sits at 69.2 percent. On Terminal-Bench 2.1, Sonnet 5 pulls 80.4 percent versus Sonnet 4.6's 67.0 percent. For multidisciplinary reasoning (Humanity's Last Exam), the model reaches 57.4 percent with tools, nearly matching Opus 4.8 at 57.9 percent. On computer use (OSWorld-Verified), Sonnet 5 posts 81.2 percent compared to 78.5 percent for its predecessor.

Ad

![](https://the-decoder.com/wp-content/uploads/2026/06/sonnet_5_benchmarks-scaled-1.webp) Sonnet 5 beats its predecessor, Sonnet 4.6, across every tested category and closes in on the pricier Opus 4.8. On knowledge work (GDPval-AA v2), Sonnet 5 even edges past Opus 4.8 with 1,618 points versus 1,615. \| Image: Anthropic

On the knowledge work benchmark GDPval-AA v2,which tests AI on real-world knowledge tasks, Sonnet 5 actually beats the larger Opus 4.8, scoring 1,618 to Opus's 1,615. Anthropic says feedback from early-access partners told the same story. Sonnet 5 acts far more agentically than previous versions, showing up in things like how it handles search tasks.

Ad

![](https://the-decoder.com/wp-content/uploads/2026/06/sonnet_5_agentic_search-scaled-1.webp) Agentic search performance on BrowseComp by effort level and cost per task. Sonnet 5 (orange) clearly outperforms Sonnet 4.6 (gray) at every level while offering cheaper entry points. Opus 4.8 (yellow) stays ahead at the highest effort settings. \| Image: Anthropic

Cybersecurity isn't a concern this time

Lately, Anthropic has been making news for models it _can't_ ship. TheUS government is blocking the company's two most capable models, Mythos 5 and Fable 5, over cybersecurity concerns. That context hangs over the Sonnet 5 launch. Anthropic is clearly eager to get ahead of any similar worries. The model wasn't trained on cybersecurity tasks, the company says, and in tests for risky capabilities like writing software exploits, it scores far below both Opus 4.8 and Mythos 5.

![](https://the-decoder.com/wp-content/uploads/2026/06/sonnet5_firefox_exploits-3840x2160-1-scaled-1.webp)Firefox 147 exploit evaluation. Like its predecessor Sonnet 4.6, Sonnet 5 couldn't develop a fully working exploit but shows a slightly higher partial control rate at 13.2 percent. Mythos 5 and Opus 4.8 are far more capable at this task. \| Image: Anthropic

Sonnet 5 does score a bit higher than its predecessor on these tasks, though. So Anthropic has switched oncyber safeguardsby default. They flag and block risky cyber usage in real time, on par with the protections already in place for Claude Opus 4.7 and 4.8. They're dialed back compared to Fable 5's guardrails, which userscomplained about almost immediately. Anthropic says it views the overall cybersecurity risk from Sonnet 5 as low.

Ad

On the safety front, the model does a better job turning down malicious requests and fending off prompt injection attacks than Sonnet 4.6, according to Anthropic. Hallucinations andsycophantic behavior, the tendency to just agree with whatever the user says, are down as well. Anthropic's full safety evaluation is in theClaude Sonnet 5 System Card.

Introductory pricing runs through August 2026

Claude Sonnet 5 is live now on all plans. It's the new default for Free and Pro users, and Max, Team, and Enterprise subscribers can access it too. Developers can plug it into Claude Code and the Claude Platform. On the API side, it goes by"claude-sonnet-5". The training cutoff is January 2026, with a one-million-token context window.

Ad

Until August 31, 2026, Anthropic is charging $2 per million input tokens and $10 per million output tokens.After that, prices jump to $3 and $15, which is what previous Sonnet models cost.

Ad

Real-world costs might tell a different story: Because the model works more agentically,it's likely to chew through more tokens per task. So even at the same per-token rate, running Sonnet 5 could end up costing more than its predecessors. The same thing happened when Opus went from 4.6 to 4.7.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now

Source:Anthropic

BETA-TEST

×

Start new search

×\|

Send

wpDiscuz

Insert

BETA-TEST

×

Start new search

×\|

Send

wpDiscuz

Insert

׋›

Region

Europe

Heat Score

89

Category

Research

Language

en