文章摘要
Claude Sonnet 5是迄今为止最具代理能力的Sonnet模型,能自主规划、使用浏览器和终端等工具,性能接近Opus 4.8但价格更低,在推理、工具使用、编码和知识工作方面较前代显著提升,且不良行为率更低。
文章总结
Anthropic 发布了 Claude Sonnet 5,这是其 Sonnet 系列中自主性最强的模型。该模型能够制定计划、使用浏览器和终端等工具,并以过去需要更大、更昂贵模型才能达到的水平自主运行。Sonnet 5 的性能接近 Opus 4.8,但价格更低,在推理、工具使用、编码和知识工作等关键自主能力上,相比前代 Sonnet 4.6 有显著提升。
安全评估显示,Sonnet 5 的不良行为率总体低于 Sonnet 4.6,在自主场景中使用更安全。同时,其执行网络安全任务的能力远低于当前的 Opus 模型。
即日起,Claude Sonnet 5 在所有计划中均可使用,并成为 Free 和 Pro 计划的默认模型。它也在 Claude Code 和 Claude 平台上提供,并推出优惠定价:每百万输入 token 2 美元,每百万输出 token 10 美元,有效期至 2026 年 8 月 31 日,之后将调整为每百万输入 token 3 美元、每百万输出 token 15 美元。
早期测试者的反馈一致认为,Sonnet 5 比前代更具自主性,能完成更复杂的任务,主动检查自身输出,且价格具有吸引力。例如,它能独立处理多步骤的软件工程工作、端到端完成销售自动化任务、以更少步骤获得相同输出质量、自主完成代码审查和修复,以及在法律研究和分析中展现出清晰的性能提升。
在安全方面,Sonnet 5 在拒绝恶意请求、抵抗提示注入攻击、减少幻觉和谄媚行为上均优于 Sonnet 4.6。虽然其在不安全行为自动审计中的总体得分更低(更安全),但相比 Opus 4.8 和 Claude Mythos Preview,其不安全行为率略高。在网络安全能力上,Sonnet 5 从未成功开发出完整的漏洞利用程序,但部分成功概率略高于 Sonnet 4.6,这可能是通用智能提升所致。因此,Anthropic 已默认启用网络安全防护措施。
评论总结
根据评论内容,主要观点和论据总结如下:
1. 性能与成本对比:Sonnet 5 性价比不如 Opus 4.8
- 多数评论认为,Sonnet 5 在中等以上推理级别时,成本与 Opus 4.8 相近但性能更差,建议直接使用更大模型。
- 关键引用:
- "Opus 4.8 beats Sonnet 5 on the pareto frontier... for certain tasks, Opus 4.8 is cheaper than Sonnet 5, and does better" (andai)
- "The cost per task chart is telling me that I should never use Sonnet 5 above medium effort level" (doctoboggan)
2. 网络安全能力显著下降
- 系统卡显示 Sonnet 5 在网络安全任务上远不如 Opus 4.8 和 Mythos 5,甚至低于前代 Sonnet 4.6,引发批评。
- 关键引用:
- "Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models" (satvikpendem)
- "American AI company status: We are now bragging about how bad our models are unironically" (jchw)
3. 定价策略与市场定位争议
- 初始定价($2/$10)被认为有竞争力,但全价时性价比不如开源模型(如 GLM 5.2)或前代产品。
- 关键引用:
- "This is much more interesting of a model at $2/$10... than at full price" (mchusma)
- "I didn't think they'd actually release a model that was worse than the open-weight frontier and at a higher price-point" (wolttam)
4. 对代理型开发的影响
- 部分用户担忧模型过度优化代理能力,反而削弱了辅助开发体验,导致转向其他模型。
- 关键引用:
- "the more models are optimized for fully agentic development, the worse they get at assisted development" (microtonal)
- "I have been moving more and more to K2.7 Code and GLM-5.2" (microtonal)
5. 整体评价:增量更新,但缺乏突破
- 多数评论认为 Sonnet 5 是“平庸的更新”,与前代相比进步有限,且存在成本上升、性能倒退等问题。
- 关键引用:
- "Seems to be another great incremental update to the workhorse" (phillipcarter)
- "Sonnet 4 was a clear step above Opus 3.x, while this is a lot muddier" (moomin)
平衡观点:少数评论认可其低推理级别的成本效率,或认为对免费用户是升级;但主流意见对性能倒退和定价策略持批评态度。