全球大模型进展新闻浏览站
中文头版English
输入关键词,快速查找已抓取新闻。

今日版 / 2026年8月12日星期三

limbo logolimbo

数据更新时间

6月30日 18:46

启用来源

17

抓取状态

真实抓取

研究进展The Decoder

Anthropic发布Claude Sonnet 5,性能逼近高端Opus系列

摘要

Anthropic推出Claude Sonnet 5模型,在各项基准测试中均超越前代Sonnet 4.6,并在GDPval-AA v2知识工作测试中以1618分略超高端Opus 4.8。公司指出,该模型在网络安全任务上的得分远低于美国政府目前封锁的模型,这可能是对当前辩论的有意信号。

背景解释

Claude Sonnet 5的发布标志着Anthropic在平衡性能与成本方面取得进展,缩小了中端与高端模型的差距。其网络安全得分较低可能意在回应监管讨论,暗示模型能力可控。对于关注AI发展的读者,这反映了模型迭代的竞争态势及安全考量。

原文译文

以下为抓取到的原文内容译文,已统一为站内阅读格式。

阅读原文

Ad

Skip to content

Anthropic's new Claude Sonnet 5 closes the gap to Opus model series

![Matthias Bastian](https://the-decoder.com/author/matthias-bastian/ "View all posts by Matthias Bastian")

Matthias Bastian[View the LinkedIn Profile of Matthias Bastian](https://www.linkedin.com/in/matthias-bastian-128b71b1/ "View the LinkedIn Profile of Matthias Bastian")

Jun 30, 2026

Image description

Anthropic

Key Points

  • Anthropic released Claude Sonnet 5, which the company calls its most agentic Sonnet yet. It can build plans on its own and use tools like browsers and terminals.
  • In benchmarks, Sonnet 5 beats its predecessor, Sonnet 4.6, across the board and closes in on the larger Opus 4.8. On real-world knowledge work tasks, it even edges past Opus 4.8.
  • The model is available now on all Anthropic platforms at an introductory discount, with pricing rising to standard Sonnet rates after August 2026.

Ask about this article…Search

Topics

Anthropic released Claude Sonnet 5. In benchmarks, it closes in on the larger Opus 4.8 and even beats it in some areas. The model is available now at an introductory price.

Anthropic calls it the most agentic Sonnet yet: it can build plans, grab tools like browsers and terminals, and work on its own at a level that just months ago only bigger, pricier models could pull off, according to the company. Sonnet 5 is meant to close that gap.

Benchmarks show a clear jump over Sonnet 4.6

Anthropic's published benchmarks show Sonnet 5 beating its predecessor Sonnet 4.6 in every tested category while gaining ground on the pricier Opus 4.8. On agentic coding, Sonnet 5 hits 63.2 percent on SWE-bench Pro, up from 58.1 percent for Sonnet 4.6. Opus 4.8 sits at 69.2 percent. On Terminal-Bench 2.1, Sonnet 5 pulls 80.4 percent versus Sonnet 4.6's 67.0 percent. For multidisciplinary reasoning (Humanity's Last Exam), the model reaches 57.4 percent with tools, nearly matching Opus 4.8 at 57.9 percent. On computer use (OSWorld-Verified), Sonnet 5 posts 81.2 percent compared to 78.5 percent for its predecessor.

Ad

![](https://the-decoder.com/wp-content/uploads/2026/06/sonnet_5_benchmarks-scaled-1.webp) Sonnet 5 beats its predecessor, Sonnet 4.6, across every tested category and closes in on the pricier Opus 4.8. On knowledge work (GDPval-AA v2), Sonnet 5 even edges past Opus 4.8 with 1,618 points versus 1,615. \| Image: Anthropic

On the knowledge work benchmark GDPval-AA v2,which tests AI on real-world knowledge tasks, Sonnet 5 actually beats the larger Opus 4.8, scoring 1,618 to Opus's 1,615. Anthropic says feedback from early-access partners told the same story. Sonnet 5 acts far more agentically than previous versions, showing up in things like how it handles search tasks.

Ad

![](https://the-decoder.com/wp-content/uploads/2026/06/sonnet_5_agentic_search-scaled-1.webp) Agentic search performance on BrowseComp by effort level and cost per task. Sonnet 5 (orange) clearly outperforms Sonnet 4.6 (gray) at every level while offering cheaper entry points. Opus 4.8 (yellow) stays ahead at the highest effort settings. \| Image: Anthropic

Cybersecurity isn't a concern this time

Lately, Anthropic has been making news for models it _can't_ ship. TheUS government is blocking the company's two most capable models, Mythos 5 and Fable 5, over cybersecurity concerns. That context hangs over the Sonnet 5 launch. Anthropic is clearly eager to get ahead of any similar worries. The model wasn't trained on cybersecurity tasks, the company says, and in tests for risky capabilities like writing software exploits, it scores far below both Opus 4.8 and Mythos 5.

![](https://the-decoder.com/wp-content/uploads/2026/06/sonnet5_firefox_exploits-3840x2160-1-scaled-1.webp)Firefox 147 exploit evaluation. Like its predecessor Sonnet 4.6, Sonnet 5 couldn't develop a fully working exploit but shows a slightly higher partial control rate at 13.2 percent. Mythos 5 and Opus 4.8 are far more capable at this task. \| Image: Anthropic

Sonnet 5 does score a bit higher than its predecessor on these tasks, though. So Anthropic has switched oncyber safeguardsby default. They flag and block risky cyber usage in real time, on par with the protections already in place for Claude Opus 4.7 and 4.8. They're dialed back compared to Fable 5's guardrails, which userscomplained about almost immediately. Anthropic says it views the overall cybersecurity risk from Sonnet 5 as low.

Ad

On the safety front, the model does a better job turning down malicious requests and fending off prompt injection attacks than Sonnet 4.6, according to Anthropic. Hallucinations andsycophantic behavior, the tendency to just agree with whatever the user says, are down as well. Anthropic's full safety evaluation is in theClaude Sonnet 5 System Card.

Introductory pricing runs through August 2026

Claude Sonnet 5 is live now on all plans. It's the new default for Free and Pro users, and Max, Team, and Enterprise subscribers can access it too. Developers can plug it into Claude Code and the Claude Platform. On the API side, it goes by"claude-sonnet-5". The training cutoff is January 2026, with a one-million-token context window.

Ad

Until August 31, 2026, Anthropic is charging $2 per million input tokens and $10 per million output tokens.After that, prices jump to $3 and $15, which is what previous Sonnet models cost.

Ad

Real-world costs might tell a different story: Because the model works more agentically,it's likely to chew through more tokens per task. So even at the same per-token rate, running Sonnet 5 could end up costing more than its predecessors. The same thing happened when Opus went from 4.6 to 4.7.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now

Source:Anthropic

BETA-TEST

×

Start new search

×\|

Send

wpDiscuz

Insert

BETA-TEST

×

Start new search

×\|

Send

wpDiscuz

Insert

׋›

来源地区

Europe

热度分

89

分类

研究进展

语言

en