Claude Opus 4.8
Description
Claude Opus 4.8 is Anthropic's upgrade to Opus 4.7 and its most capable general-access model at release, with improvements across software engineering, agentic tool use, reasoning, computer use, and knowledge-work benchmarks while shipping at the same price ($5/$25 per million input/output tokens). Performance gains include SWE-Bench Verified (88.6%), SWE-Bench Pro (69.2%), Terminal-Bench 2.1 (74.6%), GPQA Diamond (93.6%), USAMO 2026 (96.7%), Humanity's Last Exam with tools (57.9%), OSWorld-Verified (83.4%), BrowseComp (84.3% single-agent, 88.5% multi-agent), MCP-Atlas (82.2%), and GDPval-AA (1890 Elo). The alignment assessment reports honesty improvements with around a four-fold drop in letting flaws in self-written code pass unremarked, a 17-fold drop relative to Sonnet 4.6 on dishonest agentic code summaries, and broadly improved adherence to Claude's constitution. The model defaults to high effort and exposes new 'extra' (xhigh) and 'max' levels for harder problems. Launches alongside Claude Code dynamic workflows (parallel subagents that plan, execute, and verify codebase-scale migrations), effort control in claude.ai and Cowork, and a Messages API extension that accepts system entries inside the messages array so harnesses can update instructions mid-task without breaking the prompt cache. Fast mode runs at 2.5× speed at $10/$50 per million input/output tokens, three times cheaper than fast mode on previous models. Available across Claude products, the Claude API as `claude-opus-4-8`, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.
Capability Radar
Science uses a reasoning proxy when dedicated science benchmarks are unavailable.
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Agentic Capability | 7 | 72.0 | LS |
| Math Reasoning | 7 | 94.0 | LB |
| Multimodal Ranking | 13 | 63.0 | LS |
| Reasoning | 7 | 89.0 | LB |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Agents
Biology
Code
General
Grounding
Healthcare
Long Context
Math
Multimodal
AA Evaluation Indices
(Artificial Analysis)No AA evaluation data available
LLM Stats Category Scores
(LLM Stats (zeroeval))Pricing
Speed
No speed data available
Provider Price Ranking
Provider Price Ranking
16 providers
Compare pricing across different API providers for this model.