Anthropic Flagship Model

Claude Opus 4.8

Claude Opus 4.8 is Anthropic's upgraded Opus model for agentic coding, computer use, reasoning, professional knowledge work, and long-running Claude Code workflows.

May 28, 2026
official release date
69.2%
SWE-Bench Pro
83.4%
OSWorld-Verified
What Changed

Claude Opus 4.8 Capabilities

Anthropic positions Claude Opus 4.8 as a sharper, more reliable Opus release with visible gains in agent tasks, code work, computer use, and professional analysis.

Sharper Agentic Coding

Opus 4.8 raises the official SWE-Bench Pro score to 69.2% and is described by early testers as better at asking clarifying questions, catching mistakes, and resisting unsound plans.

More Reliable Tool Use

Anthropic says Opus 4.8 uses tools more cleanly, follows instructions more consistently, and is less likely than Opus 4.7 to let flaws in its own code pass unremarked.

Stronger Computer Use

The official benchmark table reports 83.4% on OSWorld-Verified, ahead of Opus 4.7, GPT-5.5, and Gemini 3.1 Pro under Anthropic's reported setup.

Effort Control

Users can choose how much effort Claude spends. High effort is the default, while extra and max effort can spend more tokens for harder tasks and long-running async workflows.

Dynamic Workflows

Claude Code can plan very large tasks, run hundreds of parallel subagents in one session, verify the output, and report back for codebase-scale migrations and similar work.

Better Honesty And Alignment

Anthropic reports stronger honesty behavior, lower misaligned behavior than Opus 4.7, and prosocial-trait highs in its alignment assessment for this release.

Official Benchmarks

Claude Opus 4.8 Benchmark Results

Anthropic's official chart compares Opus 4.8 with Opus 4.7, GPT-5.5, and Gemini 3.1 Pro across agentic coding, terminal coding, reasoning, computer use, knowledge work, and financial analysis.

Official figures include 69.2% on SWE-Bench Pro, 74.6% on Terminal-Bench 2.1, 83.4% on OSWorld-Verified, 1890 on GDPval-AA, and 53.9% on Finance Agent v2.

Official Anthropic benchmark table for Claude Opus 4.8 versus Opus 4.7, GPT-5.5, and Gemini 3.1 Pro
Claude Code

Dynamic Workflows For Large Tasks

Opus 4.8 launches with Claude Code dynamic workflows, a research-preview feature for bigger sessions that can coordinate many subagents and verify work before returning results.

Anthropic says dynamic workflows are available in Claude Code for Enterprise, Team, and Max plans, and Opus 4.8 agents can run for longer.

Official Anthropic image showing Claude Code dynamic workflows with Claude Opus 4.8
Pricing

Claude Opus 4.8 Pricing

Claude Opus 4.8 keeps regular usage pricing unchanged from Opus 4.7 and adds lower-cost fast mode pricing for higher-speed Opus use.

Regular Input
$5 / 1M tokens
Regular Output
$25 / 1M tokens
Fast Mode Input
$10 / 1M tokens
Fast Mode Output
$50 / 1M tokens

Prices are from Anthropic's May 28, 2026 Claude Opus 4.8 announcement. MyClaw plan pricing is separate from provider API token pricing.

Model Comparison

Claude Opus 4.8 vs Opus 4.7 vs GPT 5.5 vs Gemini 3.1 Pro

The official benchmark picture is clearest when Opus 4.8 is read as an incremental but broad upgrade over Opus 4.7, with particularly strong agentic coding, computer use, and knowledge-work scores.

BenchmarkOpus 4.8Opus 4.7GPT-5.5Gemini 3.1 Pro
Agentic Coding: SWE-Bench Pro69.2%64.3%58.6%54.2%
Agentic Terminal Coding: Terminal-Bench 2.174.6%66.1%78.2%70.3%
Multidisciplinary Reasoning: Humanity's Last Exam49.8% no tools / 57.9% with tools46.9% no tools / 54.7% with tools41.4% no tools / 52.2% with tools44.4% no tools / 51.4% with tools
Agentic Computer Use: OSWorld-Verified83.4%82.8%78.7%76.2%
Knowledge Work: GDPval-AA1890175317691314
Agentic Financial Analysis: Finance Agent v253.9%51.5%51.8%43.0%

Terminal-Bench 2.1 notes, OSWorld-Verified methodology updates, and additional evaluations are documented in Anthropic's Claude Opus 4.8 System Card.

MyClaw Skills

Use Claude Opus 4.8 With Long-Running Agent Skills

Pair Opus 4.8 with skills built for coding agents, GitHub issue automation, repo-scale context, terminal supervision, session history, and team updates.

Related Models

Compare Claude Opus 4.8 With Other Models

Use these pages to place Opus 4.8 in context against Claude Sonnet 5, GPT 5.6 Sol, and MiniMax M3 for agent work.

Claude Opus 4.8 FAQ

Claude Opus 4.8

Run Claude Opus 4.8 In A Hosted MyClaw Agent

Use MyClaw to pair Opus-level model access with managed runtime, files, tools, browser workflows, and persistent agent sessions.

Run Claude Opus 4.8 in MyClaw