Claude Sonnet 5
Claude Sonnet 5 is Anthropic's most agentic Sonnet model yet, built for coding, tool use, browser work, computer use, and cost-efficient autonomous agents.
What Claude Sonnet 5 Improves
Anthropic positions Claude Sonnet 5 as a lower-cost Sonnet model that narrows the gap with Opus 4.8 on agentic work while improving over Sonnet 4.6 across coding, tools, and knowledge tasks.
Agentic Planning And Tool Use
Sonnet 5 can make plans, use browsers and terminals, and run autonomously at a level that previously required larger and more expensive models.
Coding And Terminal Workflows
The system card reports 63.2 on SWE-Bench Pro, 80.4 on Terminal-Bench 2.1, and 38.8 on FrontierCode v1 for Claude Sonnet 5.
Search And Computer Use
Anthropic reports 84.7 on BrowseComp with a 10M token limit and 81.2 on OSWorld-Verified, both important signals for browser and desktop-style agents.
Cost-Performance Control
Users can adjust effort levels to trade off cost and capability, with Sonnet 5 covering a wider range of cost-performance options than Sonnet 4.6.
Safer Agentic Behavior
Anthropic says Sonnet 5 has a lower overall rate of undesirable behaviors than Sonnet 4.6 and better resistance to prompt injection in agentic contexts.
Cyber Safeguards By Default
Sonnet 5 launches with cyber safeguards enabled by default and remains much less capable than Opus 4.8 and Mythos 5 on dangerous cyber evaluations.
Official Capability Summary Across Coding And Agent Tasks
The official system card summary reports Sonnet 5 gains over Sonnet 4.6 on SWE-Bench Pro, Terminal-Bench, BrowseComp, Humanity's Last Exam, OSWorld-Verified, FrontierCode, and GDPval-AA v2.
Scores are from Anthropic's Claude Sonnet 5 System Card. BrowseComp shows single-agent and multi-agent scores. Standard Sonnet 5 results use adaptive thinking at max effort unless Anthropic notes otherwise.
Claude Sonnet 5 Scales Strongly On Agentic Search
On BrowseComp, Anthropic reports Sonnet 5 moving from 79.3 at a 1M token limit to 84.7 at a 10M token limit, while Sonnet 4.6 remains lower across the same limits.
Anthropic's launch post says the updated BrowseComp chart uses its standard methodology with a 10M token budget, context compaction, and programmatic tool calling.

Claude Sonnet 5 Reaches 81.2% On OSWorld-Verified
OSWorld-Verified measures real computer tasks in a live Ubuntu VM. Anthropic reports 81.2 for Sonnet 5, ahead of Sonnet 4.6 and close to Opus 4.8.
Anthropic notes OSWorld-Verified scores were run with updated evaluation settings that better reflect real-world performance.

FrontierCode Shows A Large Step Up From Sonnet 4.6
The FrontierCode v1 chart compares score against average cost per task across effort levels. Sonnet 5 reaches 38.8 at max effort versus 15.1 for Sonnet 4.6 in the summary table.
The chart recomputes costs at API list rates and shows effort levels from low through max. Sonnet 5 remains below Fable 5 but materially improves on Sonnet 4.6.

Automated Behavioral Audit Improves Over Sonnet 4.6 Overall
Anthropic reports that lower scores are better in its automated behavioral audit. Sonnet 5 improves over Sonnet 4.6 overall, while Opus 4.8 and Mythos Preview remain lower on some categories.
This official chart covers misaligned behavior, constitutional misalignment, and Claude Code sandbox behavior. Anthropic discusses a broader set of audit categories in Section 6.4 of the system card.

Firefox 147 Shows Low Dangerous Cyber Capability For Sonnet 5
In Anthropic's Firefox 147 exploit-development evaluation, Sonnet 5 had 0.0% working-exploit success and 13.2% any-success, far below Opus 4.8 and Mythos 5.
Anthropic says the Firefox 147 vulnerabilities were patched in Firefox 148 and that both Sonnet models had 0.0% working-exploit success.

Claude Sonnet 5 Cost And API Access
Claude Sonnet 5 launched everywhere on June 30, 2026, with introductory pricing through August 31, 2026, followed by standard Sonnet 5 pricing.
Introductory pricing runs through August 31, 2026. Anthropic notes Sonnet 5 uses an updated tokenizer, so the same input can map to roughly 1.0 to 1.35 times as many tokens depending on content type.
Claude Sonnet 5 vs. Sonnet 4.6 vs. Opus 4.8
Anthropic frames Sonnet 5 as a strict improvement over Sonnet 4.6 and a lower-cost model that can approach Opus 4.8 on selected agentic tasks at higher effort levels.
| Feature | Claude Sonnet 5 | Claude Sonnet 4.6 | Claude Opus 4.8 |
|---|---|---|---|
| Release Role | Most agentic Sonnet model yet, designed to narrow the gap with Opus-class agentic capability. | Previous best Sonnet model used as the main baseline in Anthropic's launch materials. | More generally capable Opus-class model used as the higher-capability reference. |
| API Pricing | $2 input / $10 output per 1M tokens through August 31, 2026, then $3 / $15. | Lower capability predecessor; current pricing should be checked in Anthropic's platform docs. | $5 input / $25 output per 1M tokens in Anthropic's launch-post comparison. |
| SWE-Bench Pro | 63.2 | 58.1 | Not listed in the capability summary table. |
| Terminal-Bench 2.1 | 80.4 | 67.0 | Compared in cost-performance charts; summary table compares GPT-5.5 and Gemini 3.5 Flash instead. |
| BrowseComp | 84.7 single agent and 86.6 multi agent in the summary table. | 76.2 | Cost-performance chart shows strong performance, with Sonnet 5 offering lower-cost options. |
| OSWorld-Verified | 81.2 | 78.5 | 83.4 in the OSWorld chart. |
| Cybersecurity Risk | Cyber safeguards enabled by default; lower dangerous cyber capability than current Opus models. | 0.0% working-exploit success on Firefox 147, with lower partial success than Sonnet 5. | Stronger cyber capability than Sonnet 5 and recommended by Anthropic for cybersecurity work requiring reduced guardrails. |
The supplied keyword phrase mentions Claude Sonnet 5 vs Opus 4.6; Anthropic's official launch materials compare Sonnet 5 primarily with Sonnet 4.6 and Opus 4.8.
Use Claude Sonnet 5 With Focused Agent Skills
Start Claude Sonnet 5 inside a hosted MyClaw agent, then add task-specific skills for SEO, research, content, and operational workflows instead of prompting everything from scratch.
Website SEO Skill
Pair Sonnet 5 with a repeatable website SEO audit skill so an agent can inspect pages, capture findings, and turn benchmark-level reasoning into prioritized fixes.
Open Website SEO SkillSEO Competitor Pages Skill
Use Sonnet 5 for browser-heavy competitor analysis, then let the skill organize page evidence, SERP patterns, and content gaps into a usable plan.
Open Competitor SkillSEO AEO Keyword Research Skill
Guide Sonnet 5 through keyword clustering, answer-engine opportunities, and agent-ready briefs with a focused skill instead of one-off research prompts.
Open Keyword SkillCoding Agent Skill
Use Sonnet 5 for repository search, patch planning, implementation notes, and verification steps inside a skill that keeps engineering work structured.
Open Coding SkillSession Logs Skill
Give Sonnet 5 a cleaner way to review agent history, extract decisions, and turn long-running sessions into follow-up tasks that humans can inspect.
Open Session Logs SkillSlack Skill
Connect Sonnet 5 to team workflows so a MyClaw agent can summarize updates, prepare responses, and move work across channels with clearer context.
Open Slack SkillCompare Claude Sonnet 5 With Other Agent Models
Use these pages to compare Claude Sonnet 5 against frontier reasoning, coding, and long-context models available in the MyClaw model directory.
Claude Fable 5
A Mythos-class Claude model for harder coding, vision, and long-horizon knowledge work when cost is less important than capability.
View Claude Fable 5GPT 5.6 Sol
OpenAI's flagship reasoning model for deep agentic coding, terminal tasks, biology workflows, and safeguarded cybersecurity work.
View GPT 5.6 SolMiniMax M3
A long-context multimodal coding model with benchmark coverage for terminal tasks, tool use, and agent workflows.
View MiniMax M3Claude Sonnet 5 FAQ
Model Directory
Evaluate Claude Sonnet 5 Alongside Other Agent-Ready Models
Use MyClaw to compare models in practical hosted workflows with tools, files, browser access, and managed runtime.
Run Sonnet 5 in MyClaw