Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
説明
Claude Opus 4.8 is Anthropic's upgrade to Opus 4.7 and its most capable general-access model at release, with improvements across software engineering, agentic tool use, reasoning, computer use, and knowledge-work benchmarks while shipping at the same price ($5/$25 per million input/output tokens). Performance gains include SWE-Bench Verified (88.6%), SWE-Bench Pro (69.2%), Terminal-Bench 2.1 (74.6%), GPQA Diamond (93.6%), USAMO 2026 (96.7%), Humanity's Last Exam with tools (57.9%), OSWorld-Verified (83.4%), BrowseComp (84.3% single-agent, 88.5% multi-agent), MCP-Atlas (82.2%), and GDPval-AA (1890 Elo). The alignment assessment reports honesty improvements with around a four-fold drop in letting flaws in self-written code pass unremarked, a 17-fold drop relative to Sonnet 4.6 on dishonest agentic code summaries, and broadly improved adherence to Claude's constitution. The model defaults to high effort and exposes new 'extra' (xhigh) and 'max' levels for harder problems. Launches alongside Claude Code dynamic workflows (parallel subagents that plan, execute, and verify codebase-scale migrations), effort control in claude.ai and Cowork, and a Messages API extension that accepts system entries inside the messages array so harnesses can update instructions mid-task without breaking the prompt cache. Fast mode runs at 2.5× speed at $10/$50 per million input/output tokens, three times cheaper than fast mode on previous models. Available across Claude products, the Claude API as `claude-opus-4-8`, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.
能力レーダー
ランキング
| ドメイン | #順位 | スコア | ソース |
|---|---|---|---|
| エージェント能力 | 22 | 55.0 | LS |
| コーディングランキング | 27 | 92.0 | AA |
| 総合ランキング | 27 | 87.0 | AA |
| 数学的推論 | 6 | 94.0 | LB |
| マルチモーダルランキング | 8 | 69.0 | LS |
| 推論 | 9 | 89.0 | LB |
| 科学 | 15 | 91.0 | AA |
ベンチマークスコア (LLM Stats)
(LLM Stats (zeroeval))Agents
Code
General
Healthcare
Math
Multimodal
Reasoning
Search
Vision
AA評価指数
(Artificial Analysis)LLM Statsカテゴリスコア
(LLM Stats (zeroeval))価格設定
速度
プロバイダー価格ランキング
プロバイダー価格ランキング
30 プロバイダー
このモデルの異なるAPIプロバイダー間の価格を比較。